When does OpenAI fine-tuning shut down?
Six facts from OpenAI's own pages, in date order. Inference on existing fine-tunes continues until OpenAI deprecates each base model.
| Date | What changes | Source |
|---|---|---|
| 7 May 2026 | Organisations that had never run fine-tuning can no longer create jobs. | OpenAI deprecations, checked 5 October 2026 |
| 2 July 2026 | Organisations with no fine-tuned inference in the past 60 days can no longer create jobs. | OpenAI deprecations, checked 5 October 2026 |
| 23 October 2026 | ft-gpt-3.5-turbo, ft-gpt-4, ft-gpt-4.1-nano-2025-04-14, ft-o4-mini-2025-04-16, ft-babbage-002 and ft-davinci-002 shut down. | OpenAI deprecations, checked 5 October 2026 |
| 31 October and 30 November 2026 | Existing evals become read-only, then the Evals dashboard and API shut down. OpenAI's migration note points to Promptfoo. | OpenAI deprecations, checked 5 October 2026 |
| 6 January 2027 | No customer can create new fine-tuning jobs. Existing fine-tunes keep serving until their base model is deprecated. | OpenAI deprecations, checked 5 October 2026 |
| No date | There is no documented way to export fine-tuned weights. | OpenAI fine-tuning API reference, checked 5 October 2026 |
What are the alternatives to OpenAI fine-tuning?
Overmind, Together AI, Fireworks AI and distil labs all train open-weight models you can download, and all serve them behind an OpenAI-compatible API. They differ in what you upload, how you check the new model against the old one, and how you pay. Rival cells link to that vendor's own docs, checked 5 October 2026.
| Option | What you upload | Weights download | Serving and compatibility | Evals against your current model | Open source or self-host | Price basis |
|---|---|---|---|---|---|---|
| Overmind | Traces from the Python SDK, which instruments the OpenAI SDK. Or files (CSV, TSV, JSON, JSONL, NDJSON, Parquet), or Langfuse, LangSmith, Braintrust and Galileo connectors (datasets) | Yes. Checkpoint zip from the Console or overmind model download-checkpoint (training) | OpenAI-compatible chat completions endpoint. Scales to zero; cold starts take 2 to 7 minutes (inference) | Opt-in. Score the trained model against the replies your current model returned, recorded in your traces, with five evaluator kinds (eval runs) | Yes. AGPL-3.0 platform, MIT SDK, self-host with Docker Compose (self-hosting) | Free and Pro plans plus credits. Fine-tuned inference is billed on GPU time (pricing) |
| Together AI | A prepared dataset, JSONL or Parquet in six schemas. No trace import documented (data preparation) | Yes. CLI or SDK download as .tar.zst; LoRA as merged, adapter or raw (deployment) | Dedicated endpoints only, billed per minute per replica even when idle (deployment). Chat completions documented as OpenAI-compatible (OpenAI compatibility) | A separate Evaluations product. Its compare type can test a fine-tuned endpoint against another model; you build the run (run an evaluation) | No platform source found in its GitHub org. SDKs are Apache-2.0 (GitHub) | Per 1M training tokens, $0.34 to $40 by model, with a per-job minimum. Hosting per GPU minute (pricing) |
| Fireworks AI | JSONL in OpenAI chat format. Existing OpenAI SFT files work unchanged (managed training) | Yes. firectl model download, then merge the adapter (deploying trained models) | Dedicated deployments only, billed per GPU second (serving fees). Called with the OpenAI SDK (OpenAI compatibility) | Evaluators score outputs on your data. No one-step comparison with your current model documented (evaluating trained models) | No platform source found in its GitHub org. Eval Protocol is MIT (GitHub) | Per 1M training tokens, LoRA SFT $0.50 to $10 by model size. Hosting per GPU second (pricing) |
| distil labs | Calls recorded through its OpenAI-compatible endpoint (the model must be one OpenRouter serves), exported logs as JSONL, or 20 or more examples. A teacher model then generates the training data (FAQ) | Yes, as a LoRA adapter for vLLM (local deployment). Its FAQ and pricing page differ on commercial self-hosting terms (FAQ) | Dedicated H100 endpoint behind an OpenAI-compatible API (inference) | Built in. Scored against your original production model when you start from traces (model iterations) | No open-source statement found. The Enterprise tier runs in your cloud or VPC (pricing) | Free tier with 2 training runs. Production billed per GPU hour of uptime (pricing) |
| Stay on an OpenAI base model, with prompting | Nothing to train. Examples move into the prompt | None. Fine-tuned weights have no documented export (API reference) | OpenAI's API, as today, with a new model name | Evals read-only from 31 October 2026, shut on 30 November 2026 (deprecations) | No self-host option documented (model optimization guide) | Per token (pricing) |
If all you need is the same file retrained, Fireworks is the most direct like-for-like move. Overmind fits when you want the new model trained on recent production traffic and checked against what your fine-tune actually returned. Browse the trainable bases in the model library. OpenPipe users went through the same move a vendor earlier; see Overmind vs OpenPipe.
How do you migrate a fine-tuned OpenAI model?
Start while the old fine-tune still answers, because its replies become the baseline. Each step links to the docs page that covers it.
Step 1
Export your training file
Start from the JSONL file you trained the OpenAI model on. Overmind's training contract is a messages column in OpenAI chat format. Uploads accept CSV, TSV, JSON, JSONL, NDJSON and Parquet, and land exactly as they arrived.
Upload a datasetStep 2
Trace recent traffic
Add overmind.init(providers="auto") to the Python service that calls OpenAI. It traces OpenAI SDK calls, so every request your fine-tune handles is recorded with its reply. Any OpenTelemetry exporter can send traces too.
Tracing setupStep 3
Build the eval dataset
Select recent traces, choose Add to dataset, and pick Train + eval. The evaluation share defaults to 30%. The Data Workshop's agent prepares and audits the rows, and proposes changes you approve or deny.
Datasets guideStep 4
Train an open-weight model
Run supervised fine-tuning, LoRA or full, on one of 38 open-weight bases from the Qwen, Llama, Gemma 4, LFM2.5, Antares, GPT-OSS, Nemotron and Muse Glimmer families.
Training guideStep 5
Compare before you switch
Score the trained model on the eval set. Its expected outputs are the replies your current model returned, so the check is against what runs today. The default benchmark compares the untouched base with the trained model; adding an incumbent comparison is opt-in.
The benchmarkStep 6
Switch, and keep the fallback
Point the OpenAI SDK at Overmind's endpoint and change the model id. Make live switches a capability alias with no deploy. Keep the OpenAI fine-tune callable as a fallback until OpenAI deprecates its base model.
Calling your model
The comparison uses replies recorded in your traces, so trace before the old model goes away. Serving scales to zero, so the first request after an idle spell can take 2 to 7 minutes. Code on OpenAI's Responses API moves to chat completions first.
What changes in your code?
If your code calls chat completions on the OpenAI SDK, two arguments change. Point the client at Overmind with an Overmind API key, and swap the model id.
Before, calling your OpenAI fine-tune
client = OpenAI()
reply = client.chat.completions.create(model="ft:gpt-4.1-mini-2025-04-14:your-org::abc123", messages=messages)After, calling the model you trained on Overmind
client = OpenAI(base_url="https://api.overmindlab.ai/api/v1", api_key=os.environ["OVERMIND_API_KEY"])
reply = client.chat.completions.create(model="overmind/<capability-uuid>", messages=messages)The alias overmind/<capability-uuid> follows whichever model you make live, so later swaps need no code change. To pin one version, use its ft-<job>-<base> id instead. Code built on OpenAI's Responses API moves to chat completions first. Details in the inference docs.
When should you stay on OpenAI?
- Your fine-tune is not on the 23 October list, it works, and you will not need to retrain it before 6 January 2027. It keeps serving until OpenAI deprecates its base model.
- Your job needs a proprietary GPT model, DPO, vision fine-tuning or reinforcement fine-tuning. Overmind trains open-weight models with supervised fine-tuning only.
- A current OpenAI base model with examples in the prompt matches your fine-tune on your own eval set. Then there is nothing to train.
Staying still has a deadline. After 6 January 2027 any retrain has to happen somewhere else, so plan the move before then.
How do you move off OpenAI Evals?
Existing OpenAI evals become read-only on 31 October 2026, and the Evals dashboard and API shut down on 30 November 2026. OpenAI's migration note points to Promptfoo.
In Overmind, an eval run scores one or more variants against a baseline on a frozen dataset version. Build that dataset from traces of the same traffic, so the expected outputs are what your model really returned. There are five evaluator kinds (LLM judge, deterministic, trajectory, statistical and agentic). See eval runs.
Is Overmind an alternative to OpenAI fine-tuning?
Traces
Datasets
Evals
Train
Serve
Overmind
OpenAI fine-tuning
new jobs end 6 Jan 2027
OpenAI fine-tuning as of 5 October 2026. Existing fine-tunes keep serving until their base model is deprecated; no customer can create new jobs from 6 January 2027.
Where should your fine-tune go?
Decision tree
Question 1 of 3, Which model is your fine-tune built on?
Which model is your fine-tune built on?
Four of the five answers here send you somewhere other than Overmind. Staying on OpenAI is a real option until 6 January 2027.
How does Overmind compare with OpenAI fine-tuning?
| Overmind | OpenAI fine-tuning | |
|---|---|---|
| New training jobs | Available, hosted or self-hosted | None for any customer from 6 January 2027; already closed to some organisations |
| Models | 38 open-weight bases from 230M to 72B, including Qwen, Llama, Gemma 4 and GPT-OSS | Proprietary only. SFT and DPO on gpt-4.1, gpt-4.1-mini and gpt-4.1-nano |
| Methods | Supervised fine-tuning, LoRA or full | SFT and DPO; vision on gpt-4o; reinforcement fine-tuning on o4-mini |
| Training data | Traces, file uploads, or Langfuse, LangSmith, Braintrust and Galileo connectors | Prepared JSONL, minimum 10 examples. No trace import documented |
| Weights | Downloadable checkpoint zip | No documented export |
| Serving | OpenAI-compatible chat completions; scales to zero | OpenAI's API, until the base model is deprecated |
| Evals | Eval runs with five evaluator kinds, on datasets from your traces | Read-only from 31 October 2026, shut down 30 November 2026 |
| Open source | AGPL-3.0 platform, MIT SDK, Docker Compose self-host | No self-host option documented |
| Pricing | Free and Pro plans plus credits; fine-tuned inference on GPU time | Per token. Training gpt-4.1 $25 per 1M; tuned gpt-4.1 inference $3 in, $12 out |
What does OpenAI fine-tuning do that Overmind does not?
OpenAI fine-tunes proprietary GPT models, including gpt-4.1, and offers DPO, vision fine-tuning on gpt-4o and reinforcement fine-tuning on o4-mini. Overmind trains open-weight models with supervised fine-tuning only. Existing OpenAI fine-tunes also keep serving on OpenAI's API until their base model is deprecated, with no change to your code. If yours works and you will not retrain it before 6 January 2027, staying is a reasonable choice.
What else do teams ask about moving off OpenAI fine-tuning?
Is OpenAI shutting down fine-tuning?
OpenAI is winding down self-serve fine-tuning. Since 7 May 2026 organisations that had never fine-tuned cannot start, and since 2 July neither can those without fine-tuned inference in the past 60 days. From 6 January 2027 no customer can create new jobs. Existing fine-tuned models keep serving until their base model is deprecated (OpenAI deprecations page, checked 5 October 2026).
Which fine-tuned OpenAI models shut down first?
On 23 October 2026 OpenAI removes ft-gpt-3.5-turbo, ft-gpt-4, ft-gpt-4.1-nano-2025-04-14, ft-o4-mini-2025-04-16, ft-babbage-002 and ft-davinci-002, along with several base models. Applications calling them need a tested replacement before that date; OpenAI lists recommended substitutes on its deprecations page.
Can I download the weights of my OpenAI fine-tuned model?
OpenAI documents no way to export fine-tuned weights. Moving means training a new model on the same job, starting from your original training file and your recent production traffic, then checking it against what the old model returned.
Do I have to change my code to move off OpenAI?
If your code calls chat completions, usually only the base URL, API key and model name. Overmind serves trained models on an OpenAI-compatible chat completions endpoint, so the OpenAI SDK keeps working for those calls. Code built on OpenAI's Responses API has to move to chat completions first.
What happens to my OpenAI evals?
Existing evals become read-only on 31 October 2026 and the Evals dashboard and API shut down on 30 November 2026. Overmind's eval runs score variants against a baseline with five evaluator kinds, on datasets built from your production traces.
Sources and setup guides
- OpenAI deprecations
- OpenAI supervised fine-tuning guide
- OpenAI model optimization guide
- OpenAI fine-tuning API reference
- OpenAI Evals guide
- OpenAI pricing
- Together AI fine-tuning
- Together AI pricing
- Fireworks AI managed training
- Fireworks AI pricing
- distil labs FAQ
- distil labs pricing
- Overmind datasets
- Overmind training
- Overmind inference