Skip to content

Cerebellum target: Mistral Large 3 (675B MoE, 41B active) #3

Description

@deucebucket

Model

  • Model: mistralai/Mistral-Large-3
  • Total params: 675B
  • Active params: 41B
  • Architecture: MoE with routed experts

Why Cerebellum

Massive MoE with high community demand for local deployment. Standard quants still require 48+ GB VRAM. Cerebellum's ablation-guided mixed-precision could find significant savings — expert weights in large MoE models show high variance in quantization sensitivity.

What's needed

Too large for single 24 GB GPU ablation. Need someone with multi-GPU setup to run group and layer ablation experiments.

Steps:

  1. Get BF16 GGUF + imatrix (bartowski usually has these)
  2. Group ablation — crush each tensor category to Q2_K, measure PPL
  3. Layer ablation on sacred/demotable groups
  4. Share PPL logs here

Happy to build the final override file and quant from your data.

Metadata

Metadata

Assignees

No one assigned

    Labels

    help wantedExtra attention is needed

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions