# Qwen 3 0.6B

> Canonical: https://www.overmindlab.ai/models/qwen-3-0-6b

Smallest Qwen 3 base for non-tool agents where serve cost dominates.

| Property | Value |
| --- | --- |
| Model ID | Qwen/Qwen3-0.6B |
| Provider | Qwen |
| Group | Qwen 3 |
| Parameters | 0.6B |
| Tier | Compact |
| Context window | 128K |
| Max training context | 40K |
| Tool calling | No |
| Training methods | — |
| Training cost | from $0.50 |
| Serving cost (per 1M output tokens) | from $2.50 |

## Overview

Qwen 3 gives you the widest size ladder in the catalogue: 0.6B through 32B, plus a code-specialised mixture-of-experts model. Hold the family constant, vary only size across parallel experiments, and let the benchmark pick the winner against your production model.

The widest size ladder in the catalogue, 0.6B through 32B, plus a code-specialised mixture-of-experts model. Useful when you want to hold the family constant and vary only size across parallel experiments.

## Good for

- Non-tool agents at the lowest serve footprint
- Size-ladder floor in Qwen 3 experiments
- Narrow classification or routing jobs

Train this model on your agent's data: https://docs.overmindlab.ai/models/training.md

## Related models

- [Qwen 3 4B](https://www.overmindlab.ai/models/qwen-3-4b): 4B, 128K context
- [Qwen 3 1.7B](https://www.overmindlab.ai/models/qwen-3-1-7b): 1.7B, 128K context
- [Qwen 3 8B](https://www.overmindlab.ai/models/qwen-3-8b): 8B, 128K context
- [Qwen 3 Coder 30B MoE](https://www.overmindlab.ai/models/qwen-3-coder-30b-a3b-instruct): 30B MoE (3B active), 256K context