AIAT·V2 is a thesis experiment (Philosophy & Artificial Intelligence, Sapienza University of Rome) that gives four LLM agents $1,000 each and lets them trade crypto perpetual futures autonomously on the Hyperliquid testnet — a decision every 15 minutes, guardrails, stop-losses, and a fully persisted audit trail.
| Model | Provider | Tier | |
|---|---|---|---|
| 🟡 | Claude Opus 4.8 | Anthropic | USA · premium |
| 🟢 | GPT 4.1 mini | OpenAI | USA · cheap |
| 🟣 | Qwen3.7-Max | Alibaba | CN · premium |
| 🔵 | DeepSeek V4 Flash | DeepSeek | CN · cheap |
Not a trading-bot showcase — a comparative behavioral study. The four models get the identical prompt, the identical market context, and the identical capital; what differs is the mind. The experiment measures behavioral differences between providers — trade frequency, risk appetite, reasoning style, reliability (schema compliance, fallbacks), and the true cost of intelligence — across two pre-registered axes: USA vs CN and premium vs cheap. It is not about "who gets rich": it's testnet money, and the interesting output is the audit trail, not the PnL.
Hyperliquid testnet
⇅ orders · fills ⇅ market data · account state
┌───────────────────────────┐ ┌────────────────────────────────────┐
│ 4× agent-<provider> │ │ context-orchestrator │
│ one LLM + one wallet each│ │ one context_snapshot per 15-min │
│ guardrails → execution │ │ tick + ClosureReconciler (books │
│ │ │ SL/TP closures hit between ticks) │
└─────────────┬─────────────┘ └─────────────────┬──────────────────┘
│ decisions · positions · costs │ snapshots · closures
▼ ▼
PostgreSQL 16 (Railway) — the single audit trail
│ reads (SELECT-only role)
▼
AIAT·V2 Dashboard (separate repo)
- One codebase, six services — the same image runs as orchestrator or agent via
AIAT_SERVICE_ROLE; Postgres is the sixth Railway service - Market-context parity — only the orchestrator talks to external sources; the four
agents read the same frozen
context_snapshot, nobody fetches anything mid-run - Guardrails before the exchange — leverage/size clamps, forced HOLD, testnet-only
startup check,
Decimaleverywhere money flows - ClosureReconciler — stop-loss/take-profit closures that fire between ticks are detected on-chain and booked at orchestrator level, keeping DB and chain in agreement
- Cross-model isolation — every agent query is filtered by
model_id, enforced by end-to-end tests with a DB trap
- Certified prompt — all four models receive the same prompt template; its hash is persisted with every run and frozen for the whole experiment
- Frozen blueprint + ADR governance — the PRD is tagged
prd-v2-frozen; every deviation or evolutive decision is an Architecture Decision Record indocs/decisions/(34 accepted so far) - Pre-registered baselines — cash, buy & hold, and EMA-momentum curves are declared in
docs/RESEARCH_DESIGN.mdbefore the run, and computed from the same context snapshots the models see - DB ↔ chain reconciliation — positions are verified fill-by-fill against on-chain
userFills; divergences are detected, root-caused, and documented — see the M6.1 methodological note for full transparency on the shakedown run
src/aiat/ one Python package, five service roles (AIAT_SERVICE_ROLE dispatch):
domain · context · llm · execution · orchestration · baselines ·
db · config · observability · prompts
alembic/ schema migrations — the database is never edited by hand
docs/ PRD_V2 (frozen blueprint) · RESEARCH_DESIGN · M6.1 methodological note ·
decisions/ (ADRs) · runbooks
scripts/ one-shot ops: experiment seed, fee backfill, audited data repairs, baselines
tests/ 808 tests: unit · integration · e2e (isolation, invariants) · VCR cassettes
tools/ gate_check.sh — milestone gate runner
docker/ multi-stage Dockerfile: one image, role picked via env
legacy/ the V1 prototype, preserved
asset/ README media (mascot)
The experiment ships with a live observability layer — equity race with baselines, decision ledger, side-by-side reasoning, cost of intelligence, reliability — built as a separate read-only repo:
- Repo: aithos-rr/AI-Agent-for-Trading-Dashboard
- Live: https://dashboard-production-898d.up.railway.app/
Comparative study of the decision quality of 4 LLMs (USA / CN × premium / cheap) as autonomous trading agents: same prompt, same capital, same market — different minds.
The full research design (3 research questions, hypotheses, baselines) lives in
docs/RESEARCH_DESIGN.md; the technical blueprint in
docs/PRD_V2.md.
Aithos · Sapienza Università di Roma · Hyperliquid testnet · no real money involved