Back to News & insightsAI models

Mixture of experts: what active parameters actually tell you

Read sparse-model specifications with a clearer view of routing, memory, serving costs, and the evidence needed to choose a deployment.

Editorial guide · Updated September 27, 2026 · 4 min read
A silver routing hub connects to separate glass chambers, with only a subset illuminated.

A model specification can list a large total parameter count beside a much smaller active count. Both numbers can be correct. They describe different aspects of the system, and neither is a shortcut for predicting how quickly a model will answer your question.

Mixture-of-experts architectures make this distinction especially important. For someone choosing a model, the useful questions are what computation runs for each token, where the weights live, and how the complete service behaves under the workload it will actually receive.

Experts are learned components, not named employees

In a sparse mixture-of-experts layer, a routing mechanism chooses a subset of expert components to process a token. The Mixtral paper provides a concrete historical example: its router selects two of eight feedforward experts at each layer. That is an architectural description of that model, not a rule for every sparse model.

The word expert is easy to overinterpret. It does not imply that one component is a lawyer, another is a programmer, and a third is a translator. Learned specialization need not match human categories. A model can also route different tokens in the same sentence to different experts.

Separate capacity, computation, and residency

Total parameters describe the model's stored capacity. Active parameters describe the subset participating in a particular computation under the model's accounting convention. Serving memory also includes temporary activations, attention state, runtime overhead, and any extra copies required by the deployment.

An inactive expert still needs to be available when a later token selects it. Depending on the serving design, its weights may remain on an accelerator, live elsewhere, or be transferred as needed. Consequently, a small active count does not establish that the full model fits in a small device's memory.

When two model cards use different counting conventions, compare the architecture notes before comparing the numbers. Shared layers, embedding weights, and routing components can complicate a simple ratio. Treat an unlabeled parameter number as incomplete information.

A procurement example with three separate constraints

Imagine a research team building a multilingual document assistant. This is an illustrative scenario. The team has enough hardware to store a sparse model, but many researchers submit long documents at the same time. Its first bottleneck may be request memory and scheduling rather than the arithmetic for one token.

The team should write down three requirements: an acceptable answer-quality threshold, a maximum user wait under expected concurrency, and an operating envelope the hardware can sustain. Passing a single-user speed test satisfies only a small part of that contract.

Create separate test groups for short questions, long-document questions, and mixed traffic. Keep document sizes and output limits representative. Record queue time as well as generation time, because the experience begins when a person submits a request, not when an accelerator becomes available.

Routing has operational consequences

If a deployment distributes experts across devices, moving intermediate data between them can contribute to latency. The impact depends on the architecture, interconnect, implementation, and traffic. A public description of sparse computation does not reveal the performance of a particular hosting setup.

Ask a hosting vendor for measurements with a stated model revision, precision, context distribution, and concurrency. If those details are unavailable, run a bounded trial. Observe variation across requests instead of presenting the best response as typical performance.

For self-hosting, inspect device utilization and memory pressure during the same test. Do not attribute every slowdown to routing: tokenization, retrieval, and application queues can also dominate. Profiling should narrow the explanation before configuration changes begin.

Compare completed work

The document assistant should be judged on whether answers preserve the source meaning, identify missing evidence, and cite the correct passage. Sparse and dense models can both fail those requirements. Architecture is a useful diagnostic lens, but the product still needs a behavioral evaluation.

Include requests in each supported language and documents with ambiguous terminology. Keep failures visible. A low operating cost per generated token is less useful if the system repeatedly needs another attempt or sends many answers for manual correction.

What to put on the decision sheet

Record total and active parameters separately, then add observed memory, quality, latency, concurrency, and the exact serving configuration. Mark unreported values as unknown. Include the applicable model licence and the practical work needed to maintain the deployment.

Choose the system that meets the task within those constraints. That decision might favor a sparse model, a dense model, or a hosted service whose internal architecture is not disclosed. A defensible comparison acknowledges missing information instead of turning a parameter count into an imagined performance score.

Research background

Mixtral of Experts documents a specific sparse architecture and its experiments. The operational checklist and document-assistant scenario above are original application guidance, not results reported in that paper.

Read the original research paper on arXiv

An original editorial guide. Provider capabilities and documentation can change. Follow the linked sources and test the exact model or service before relying on it.