Superseded by Neo-V3. Archived for history.
This repository preserves the control plane layer of NEO-LAB: the deterministic controller (neo_loop.ps1) and chat client (neo_chat.ps1) that classify intent, load explicit state and personas, route each request to exactly one model, and enforce the output contracts. It sits on top of the atomic file-based queue substrate and is where routing rules, modes, and governance constraints live.
NEO-LAB as a whole is a local, deterministic cognitive control plane for orchestrating multiple locally hosted large language models (LLMs) as role-separated inference engines under explicit governance, bounded memory, and fully inspectable state.
All execution is local. All state is explicit. All behavior is deterministic.
NEO-LAB implements a control-plane architecture for AI inference rather than a traditional chatbot interface. It treats LLMs as stateless, interchangeable compute engines governed by a stateful, deterministic controller.
The system prioritizes:
- Predictability over emergence
- Inspectability over opacity
- Governance over autonomy
This makes NEO-LAB suitable for professional, technical, and safety-conscious workflows.
NEO-LAB is:
- A deterministic router (rule-based intent → model selection)
- A governed control plane (explicit state, modes, and contracts)
- A file-driven IPC system (atomic, concurrency-safe)
- A local-first AI system (offline-capable, no cloud dependency)
- A single-response generator for summaries, detailed explanations, and multi-page reports
NEO-LAB is not:
- An autonomous agent
- A self-modifying system
- A tool-execution framework
- A blended or ensemble model
- A cloud service
NEO-LAB uses an atomic per-message file queue to eliminate race conditions and ensure deterministic execution.
queue_v2/
├─ inbox/ # One JSON file per user message
├─ outbox/
│ └─ <message_id>/
│ ├─ status.json # Execution phase, progress counters
│ └─ response.txt # Streaming model output
└─ processed/ # Archived input messages
User
↓
neo_chat.ps1
↓ (atomic JSON message)
queue_v2/inbox/<id>.json
↓
neo_loop.ps1
├─ intent classification
├─ explicit state & persona loading
├─ deterministic model routing
├─ streaming inference
└─ bounded memory update
↓
queue_v2/outbox/<id>/response.txt
Key properties:
- No message overwrites
- Safe concurrency
- Replayable execution
- Fully inspectable artifacts
NEO-LAB routes requests deterministically. Exactly one model is active per request.
Typical role mapping (configurable):
| Cognitive Role | Model |
|---|---|
| Chat / Persona | dolphin-llama3 |
| Code | deepseek-coder-v2 |
| Analysis | deepseek-r1 |
| Vision | qwen2.5-vl |
Model availability is queried dynamically via Ollama. If a target model is unavailable, NEO-LAB fails closed or falls back safely.
NEO-LAB is designed to produce one complete response per request, regardless of domain.
Supported one-shot modes:
/summary <topic>— concise, complete response/detail <topic>— detailed technical explanation/report <topic>— single multi-page structured report
When enabled (/reportmode on), NEO-LAB enforces a strict output contract.
Required sections:
- Executive Summary
- Scope & Assumptions
- Core Analysis
- Methods / Models / Math (if applicable)
- Implementation (if applicable)
- Risks, Limitations, Verification Checklist
- References / Source Guidance (if applicable)
No follow-up questions. No partial answers. One complete professional deliverable.
NEO-LAB streams output incrementally while models generate.
Progress is written to:
queue_v2/outbox/<id>/status.json
Including:
- Current execution phase
- Characters written
- Streaming activity indicators
This ensures transparency during long-running analyses and reports.
Memory is:
- JSON-based
- Explicitly bounded
- Stored on disk
- Separated by intent
queue/
├─ memory_chat.json
├─ memory_code.json
├─ memory_analysis.json
└─ memory_vision.json
There are:
- No embeddings
- No hidden vectors
- No implicit recall
Memory can be inspected or cleared at any time.
NEO-LAB is designed with explicit safety constraints:
- No autonomy
- No self-execution
- No self-modification
- No privilege escalation
- Human remains root authority
The system favors control, auditability, and predictability over emergent behavior.
- Windows 10 / 11
- PowerShell 5.1
- Ollama (local inference server)
Recommended Ollama models:
- dolphin-llama3
- deepseek-coder-v2
- deepseek-r1
- qwen2.5-vl
Start the control loop (Terminal A):
cd <repo-root>
.\neo_loop.ps1Start the chat client (Terminal B):
cd <repo-root>
.\neo_chat.ps1Example:
/report Write a multi-page engineering report explaining why atomic IPC queues prevent race conditions.
A simple QA harness validates report structure and completeness:
powershell -ExecutionPolicy Bypass -File tests\qa\qa_harness.ps1Results are written to tests/qa/qa_results.json.
See LICENSE.
NEO-LAB provides general informational output only. For legal, medical, or financial decisions, consult qualified professionals and primary sources.