# Qwen 3.5 4B

> Canonical: https://www.overmindlab.ai/models/qwen-3-5-4b

Top of the Qwen 3.5 compact band, still trainable at 256K context.

| Property | Value |
| --- | --- |
| Model ID | Qwen/Qwen3.5-4B |
| Provider | Qwen |
| Group | Qwen 3.5 |
| Parameters | 4B |
| Tier | Compact |
| Context window | 256K |
| Max training context | 256K |
| Tool calling | Yes |
| Training methods | — |
| Training cost | from $0.70 |
| Serving cost (per 1M output tokens) | from $4.00 |

## Overview

Qwen 3.5 is the default recommendation for most agents on Overmind. Every size shares a 256K inference context, tool calling starts at 0.8B, and the compact models train at their full 256K context, the longest trainable window in the catalogue.

The default recommendation for most agents. A 256K context window at every size, tool calling from 0.8B up, and compact models trainable at their full 256K context, the longest in the catalogue.

## Good for

- General-purpose agents that call tools
- Long-trace training on compact bases
- Size ladders from 0.8B through a 35B MoE

Train this model on your agent's data: https://docs.overmindlab.ai/models/training.md

## Related models

- [Qwen 3.5 2B](https://www.overmindlab.ai/models/qwen-3-5-2b): 2B, 256K context
- [Qwen 3.5 0.8B](https://www.overmindlab.ai/models/qwen-3-5-0-8b): 0.8B, 256K context
- [Qwen 3.5 9B](https://www.overmindlab.ai/models/qwen-3-5-9b): 9B, 256K context
- [Qwen 3.5 27B](https://www.overmindlab.ai/models/qwen-3-5-27b): 27B, 256K context