MoE COMPRESSION / QUALITY REPAIR

Shrink Large MoE.
Fix the Quality Drop.

Compression degrades quality. We repair what it costs and deliver the result.

Compression can reduce the number of GPUs required to run large MoE models. GPU-count estimates are based on weight-resident capacity; the production count including KV cache is determined during the pre-deployment assessment.
470.2GB156.1GB
Qwen3-235B-A22B / measured bytes / Apache-2.0
~69%
Recovered about 69% of factuality lost in AWQ compression of Qwen3-30B Target Qwen3-30B-A3B / Compression base AWQ-style int4 / Evaluation TruthfulQA / Metric Ragas FactualCorrectness / Judge gpt-4o
WHAT LINE A DOES

Make large MoE models smaller, then fix quality lost in compression.

Compression and quality repair are handled separately so the model can remain small after repair.

01

Compress

Reduce the size of the MoE model.

02

Fix quality

Repair quality lost during compression.

SERVICES

Order each stage separately.

You may stop after measurement. There is no obligation to continue to diagnosis or repair.

AvailableQualityAudit External

Measure

Measure a compressed model you already hold against the original under the same conditions.

You provide
The locations of the compressed model you already hold and its original
You receive
Measured differences from the original with confidence intervals, a full record of measurement conditions, and limitations

We do not compress or repair the model in this SKU, and we do not perform diagnostic processing.

QualityAudit Deep

Diagnose

Create an assessment before deciding whether to proceed to repair.

You receive
An assessment of how far the quality can be repaired
QualityFix

Fix

Repair quality for a model that falls within the current supported scope.

You receive
The repaired model and Quality Certificate
PROCESS

Process

Scope and measurement conditions are set before measurement begins.

  1. 0130-minute online
    discussion
  2. 02Use the pre-check questionnaire
    to confirm the model is in scope
  3. 03Set measurement conditions and
    the acceptance value before measurement
  4. 04Measure
RESULTS

Measured results

Only measured byte sizes and results produced under fixed evaluation conditions are shown here.

RakutenAI-2.0-8x7B

93.7GB29.8GB

Measured bytes / Apache-2.0

LLM-jp-3.1-8x13B

146.3GB49.5GB

Measured bytes / Apache-2.0

Qwen3-235B-A22B

470.2GB156.1GB

Measured bytes / Apache-2.0

QUALITY REPAIR69%

Recovered about 69% of factuality lost in AWQ compression of Qwen3-30B

Target Qwen3-30B-A3B / Compression base AWQ-style int4 / Evaluation TruthfulQA / Metric Ragas FactualCorrectness / Judge gpt-4o

SCOPE

Current scope

We state the current supported range as-is.

MoE onlyCurrent work is limited to MoE models.
classic AWQ int4 onlyThe current compression base is limited to this format.
bf16 original requiredAn original model is required for comparison.
Customer-owned GPUExecution takes place in the customer's GPU environment.

These are records of measurements we performed ourselves on publicly available models. They are published as samples showing the record format. No customer work is included.

View public records →
CONTACT

We first confirm the target model.

We start with a 30-minute online discussion and then confirm scope with a pre-check questionnaire.

GPU environment