The repository is ~450 KB. Everything large is regenerated on the target: engine sources, Docker images, datasets and results are all gitignored.
The framework is standalone. It does not need the AliSQL/, server/ or
ann-benchmarks/ checkouts that sit beside it on the original machine. When
they are absent, prepare-sources.sh and prepare-harness.sh clone from
upstream instead. Copying them across is an optimisation, not a requirement.
# Debian / Ubuntu
sudo apt-get update
sudo apt-get install -y docker.io python3-yaml git curl
sudo usermod -aG docker "$USER" # then log out and back inCheck before starting:
| Resource | Needed | Why |
|---|---|---|
| Disk | ~80 GB | ~10 GB images, ~25 GB engine sources, ~5 GB datasets, the rest index data |
| RAM | 32 GB+ | the normalized pass defaults to a 16 GB server limit |
| Cores | 8+ physical | server and client containers are pinned to disjoint CPU sets |
| Network | first run only | cloning sources and downloading datasets |
df -h / # want 80 GB+ free
nproc; free -g
docker info >/dev/null && echo "docker ok"
python3 -c 'import yaml' && echo "pyyaml ok"# on the source machine
cd ~/AI_WORK/VECTOR_RESEARCH/vector-bench
git remote add origin <your-repo-url>
git push -u origin master
# on the target
git clone <your-repo-url> vector-bench
cd vector-bench# from the source machine
rsync -avz --exclude-from=<(printf 'sources/\nwork/\ndatasets/\nresults/\n.git/\n') \
~/AI_WORK/VECTOR_RESEARCH/vector-bench/ user@target:~/vector-bench/Either way, confirm the scripts survived the trip with their permissions:
chmod +x run-benchmark.sh scripts/*.sh tests/*.shcd ~/vector-bench
# If the build host IS the benchmark host, use native so each engine gets the
# best SIMD path its CPU offers (including AVX-512, which this framework's
# original host did not have). Otherwise keep x86-64-v3 for portability.
./run-benchmark.sh build --march nativeExpect roughly:
| Engine | Cold build |
|---|---|
| pgvector | ~10 min |
| MariaDB | ~40 min |
| AliSQL | 1.5–3 h |
AliSQL dominates because its CMake compiles the bundled DuckDB unconditionally
on Linux and redirects that output to /dev/null — the build looks stalled for
a long stretch. It is not. Build one engine at a time if you want to start
measuring the other two sooner:
./run-benchmark.sh build --engines pgvector --march native
./run-benchmark.sh build --engines mariadb --march native
./run-benchmark.sh build --engines alisql --march native # start this earlyNever mix
-marchbetween engines. Rebuilding one with a different value turns the benchmark into a comparison of compiler flags. The value is baked into each image and reported in the manifest, so a mismatch is visible after the fact — but it wastes the run.
./run-benchmark.sh fetch --list # sizes and roles
./run-benchmark.sh fetch # the four profile datasets, ~5 GBDownloads are resumable and size-verified, so an interrupted transfer will not be mistaken for a complete dataset.
python3 -m pytest tests/ -q # 73 unit tests, ~2 s
./tests/verify-alisql-traps.sh # 8 engine-behaviour checks
./run-benchmark.sh run --profile smoke # ~30 min, all three enginesDo not skip the smoke profile. It exercises every stage end to end and is far cheaper than discovering a broken image eight hours into a full run.
There is also a one-minute synthetic cycle that needs no dataset download:
./tests/make-tiny-dataset.sh
./run-benchmark.sh run --profile dev./run-benchmark.sh run --profile mainmain is the profile sized to actually be run — two datasets at 1M scale,
roughly two days of ingest. full --resource-pass both describes the complete
measurement space and is about eleven days on this hardware; see
07-planning-a-run.md before choosing it.
Give the machine to it. Nothing else running — no CI, no other containers, no builds. Concurrent load distorts results by around 2× and the harness cannot detect it: the Validity section reports CPU, SIMD and cpuset problems, but a competing workload is invisible to it.
Resumable after an interruption:
./run-benchmark.sh run --profile main --resume --run-id main-<timestamp>After the first run, confirm the target is what you think it is:
python3 -c "
import json; m=json.load(open('results/<run-id>/run-manifest.json'))
c=m['host']['cpu']
print('cpu ', c['model'])
print('avx512 ', c['has_avx512'])
print('hybrid ', c['hybrid'])
print('cpuset ', m['config']['resolved_resources']['server_cpuset'])
print('warnings '); [print(' -', w) for w in m['warnings']]
"has_avx512: true is the main thing worth confirming — both MariaDB MHNSW and
AliSQL VIDX document AVX-512 distance kernels, and results from a machine
without it do not transfer to one with it.
rsync -avz user@target:~/vector-bench/results/<run-id>/ ./results/<run-id>/The run directory is self-contained: manifest, raw records, charts and both
report formats. report.html inlines its charts and needs no network.
To regenerate the report locally from copied results:
./run-benchmark.sh report --run-dir results/<run-id>Most changes do not require rebuilding images. The images contain only the compiled servers plus a pinned Python stack; the orchestrator, ops harness, report generator, profiles and resource configs are all read from the working tree at run time.
cd ~/vector-bench
git pull
python3 -m pytest tests/ -q # confirm the pull is saneThat is usually the whole procedure. Decide whether more is needed by what the pull touched:
| Changed path | What to do |
|---|---|
orchestrator/, harness/, report/, config/profiles/, config/resources/, docs/ |
Nothing. Picked up on the next run. |
overlay/ |
Nothing — the working copy is refreshed on every run. |
docker/, config/engines/ (build flags or source tag) |
Rebuild that engine: ./run-benchmark.sh build --engines <name> --march native |
To check whether a rebuild is required after pulling:
git diff --name-only HEAD@{1} HEAD -- docker/ config/engines/Empty output means your images are still valid.
Re-running the report over results you already have needs no new measurement at all — useful after a report-generator fix:
./run-benchmark.sh report --run-dir results/<run-id>If the target already has MariaDB, AliSQL or ann-benchmarks checkouts, point at them and skip the upstream clones:
export VB_REPO_MARIADB=/path/to/server
export VB_REPO_ALISQL=/path/to/AliSQL
export VB_REPO_ANNB=/path/to/ann-benchmarks
./run-benchmark.sh buildThey are read only — the framework clones or git archives out of them and
never writes to them.