- Six-route literature and reproduction structure established.
- Route 1 official LM-Polygraph estimator pilot completed.
- Prompt-level stress labels separated from manually audited generation labels.
Next Route 1 milestones:
- Replace the diagnostic prompt set with claim-level factuality benchmarks.
- Add sampling-based semantic entropy and semantic-volume comparisons.
- Evaluate several text-only open-weight model families and decoding settings.
- Report calibration, compute cost, and risk--coverage with uncertainty intervals.
- Bayesian regression benchmark.
- Probabilistic scoring.
- MCMC diagnostics.
- Repeated-split comparison.
- Part I reframed as the Bayesian predictive foundation for the doctoral programme.
- README expanded into a linked multi-level research map.
- Part III and Part IV documentation scaffolds.
Follow-up:
- Audit Part I sampler and written prior consistency in a dedicated PR before any paper-style release.
Milestones:
- Part II documentation scaffold.
- Text-only hallucination-risk dataset prototype.
- Feature extraction for uncertainty, consistency, and evidence.
- Bayesian logistic risk model.
- Calibration metrics: NLL, Brier, ECE, AUROC, AUPRC.
- Risk-coverage analysis.
- Explicit observation models for noisy labels and correlated evidence signals.
Milestones:
- Choose multimodal hallucination benchmarks.
- Define hallucination types.
- Extract visual grounding / evidence features.
- Fit hierarchical Bayesian risk models.
- Evaluate by hallucination type and dataset.
Milestones:
- Define actions: answer, abstain, verify, regenerate.
- Define decision costs.
- Compare posterior-mean thresholds with credible-upper-bound thresholds.
- Evaluate selective risk and coverage.
- Study practical reliability trade-offs.