Skip to main content

Run a Compression Job

Shrink a model's size with CompactifAI while healing back quality.

Before you start

The model you want to compress must already be in your project's Model HUB catalog. If it isn't there yet, import Qwen/Qwen3-0.6B from Hugging Face first, see Upload a Model from Hugging Face.

Steps

  1. Open Compression Jobs — Go to Product Guide → Compression Jobs.
  2. Click New Compression Job — This opens CompactifAI Studio.
  3. Choose a Model preset — Pick the model from Model HUB you want to compress.
  4. Choose a Use Case — e.g. Agentic (Tool calling) — this tunes the compression strategy to how the model will actually be used.
  5. (Optional) Name the job — Leave blank to auto-generate a unique name.
  6. Set the Compression rate — Drag the slider toward your target size reduction; stay in the marked "best range" for the most reliable results.
  7. Pick Healing steps — Choose how many recovery stages to run, from 1 (Recovery only) up to 4 (Recovery, Consolidation, Adaptation, and Localization); each stage you pick includes every stage before it. See Healing Steps for what each one does.
  8. Choose Compression speed — e.g. Fast (16 GPUs). Faster settings finish sooner but use more GPU capacity at once; check the estimated duration shown.
  9. Review Estimated cost — Confirm the live cost range shown before launching.
  10. Click Compress — The job appears at the top of the Recent Workflow list and runs through to Succeeded, Failed, or Cancelled.
  11. Track or cancel it — Click View on the job row for details, or Cancel while it's still running.