Skip to main content

When does OpenAI fine-tuning shut down?

Six facts from OpenAI's own pages, in date order. Inference on existing fine-tunes continues until OpenAI deprecates each base model.

OpenAI fine-tuning wind-down dates, checked 5 October 2026
DateWhat changesSource
7 May 2026Organisations that had never run fine-tuning can no longer create jobs.OpenAI deprecations, checked 5 October 2026
2 July 2026Organisations with no fine-tuned inference in the past 60 days can no longer create jobs.OpenAI deprecations, checked 5 October 2026
23 October 2026ft-gpt-3.5-turbo, ft-gpt-4, ft-gpt-4.1-nano-2025-04-14, ft-o4-mini-2025-04-16, ft-babbage-002 and ft-davinci-002 shut down.OpenAI deprecations, checked 5 October 2026
31 October and 30 November 2026Existing evals become read-only, then the Evals dashboard and API shut down. OpenAI's migration note points to Promptfoo.OpenAI deprecations, checked 5 October 2026
6 January 2027No customer can create new fine-tuning jobs. Existing fine-tunes keep serving until their base model is deprecated.OpenAI deprecations, checked 5 October 2026
No dateThere is no documented way to export fine-tuned weights.OpenAI fine-tuning API reference, checked 5 October 2026

What are the alternatives to OpenAI fine-tuning?

Overmind, Together AI, Fireworks AI and distil labs all train open-weight models you can download, and all serve them behind an OpenAI-compatible API. They differ in what you upload, how you check the new model against the old one, and how you pay. Rival cells link to that vendor's own docs, checked 5 October 2026.

Alternatives to OpenAI fine-tuning compared, checked 5 October 2026
OptionWhat you uploadWeights downloadServing and compatibilityEvals against your current modelOpen source or self-hostPrice basis
OvermindTraces from the Python SDK, which instruments the OpenAI SDK. Or files (CSV, TSV, JSON, JSONL, NDJSON, Parquet), or Langfuse, LangSmith, Braintrust and Galileo connectors (datasets)Yes. Checkpoint zip from the Console or overmind model download-checkpoint (training)OpenAI-compatible chat completions endpoint. Scales to zero; cold starts take 2 to 7 minutes (inference)Opt-in. Score the trained model against the replies your current model returned, recorded in your traces, with five evaluator kinds (eval runs)Yes. AGPL-3.0 platform, MIT SDK, self-host with Docker Compose (self-hosting)Free and Pro plans plus credits. Fine-tuned inference is billed on GPU time (pricing)
Together AIA prepared dataset, JSONL or Parquet in six schemas. No trace import documented (data preparation)Yes. CLI or SDK download as .tar.zst; LoRA as merged, adapter or raw (deployment)Dedicated endpoints only, billed per minute per replica even when idle (deployment). Chat completions documented as OpenAI-compatible (OpenAI compatibility)A separate Evaluations product. Its compare type can test a fine-tuned endpoint against another model; you build the run (run an evaluation)No platform source found in its GitHub org. SDKs are Apache-2.0 (GitHub)Per 1M training tokens, $0.34 to $40 by model, with a per-job minimum. Hosting per GPU minute (pricing)
Fireworks AIJSONL in OpenAI chat format. Existing OpenAI SFT files work unchanged (managed training)Yes. firectl model download, then merge the adapter (deploying trained models)Dedicated deployments only, billed per GPU second (serving fees). Called with the OpenAI SDK (OpenAI compatibility)Evaluators score outputs on your data. No one-step comparison with your current model documented (evaluating trained models)No platform source found in its GitHub org. Eval Protocol is MIT (GitHub)Per 1M training tokens, LoRA SFT $0.50 to $10 by model size. Hosting per GPU second (pricing)
distil labsCalls recorded through its OpenAI-compatible endpoint (the model must be one OpenRouter serves), exported logs as JSONL, or 20 or more examples. A teacher model then generates the training data (FAQ)Yes, as a LoRA adapter for vLLM (local deployment). Its FAQ and pricing page differ on commercial self-hosting terms (FAQ)Dedicated H100 endpoint behind an OpenAI-compatible API (inference)Built in. Scored against your original production model when you start from traces (model iterations)No open-source statement found. The Enterprise tier runs in your cloud or VPC (pricing)Free tier with 2 training runs. Production billed per GPU hour of uptime (pricing)
Stay on an OpenAI base model, with promptingNothing to train. Examples move into the promptNone. Fine-tuned weights have no documented export (API reference)OpenAI's API, as today, with a new model nameEvals read-only from 31 October 2026, shut on 30 November 2026 (deprecations)No self-host option documented (model optimization guide)Per token (pricing)

If all you need is the same file retrained, Fireworks is the most direct like-for-like move. Overmind fits when you want the new model trained on recent production traffic and checked against what your fine-tune actually returned. Browse the trainable bases in the model library. OpenPipe users went through the same move a vendor earlier; see Overmind vs OpenPipe.

How do you migrate a fine-tuned OpenAI model?

Start while the old fine-tune still answers, because its replies become the baseline. Each step links to the docs page that covers it.

  1. Step 1

    Export your training file

    Start from the JSONL file you trained the OpenAI model on. Overmind's training contract is a messages column in OpenAI chat format. Uploads accept CSV, TSV, JSON, JSONL, NDJSON and Parquet, and land exactly as they arrived.

    Upload a dataset
  2. Step 2

    Trace recent traffic

    Add overmind.init(providers="auto") to the Python service that calls OpenAI. It traces OpenAI SDK calls, so every request your fine-tune handles is recorded with its reply. Any OpenTelemetry exporter can send traces too.

    Tracing setup
  3. Step 3

    Build the eval dataset

    Select recent traces, choose Add to dataset, and pick Train + eval. The evaluation share defaults to 30%. The Data Workshop's agent prepares and audits the rows, and proposes changes you approve or deny.

    Datasets guide
  4. Step 4

    Train an open-weight model

    Run supervised fine-tuning, LoRA or full, on one of 38 open-weight bases from the Qwen, Llama, Gemma 4, LFM2.5, Antares, GPT-OSS, Nemotron and Muse Glimmer families.

    Training guide
  5. Step 5

    Compare before you switch

    Score the trained model on the eval set. Its expected outputs are the replies your current model returned, so the check is against what runs today. The default benchmark compares the untouched base with the trained model; adding an incumbent comparison is opt-in.

    The benchmark
  6. Step 6

    Switch, and keep the fallback

    Point the OpenAI SDK at Overmind's endpoint and change the model id. Make live switches a capability alias with no deploy. Keep the OpenAI fine-tune callable as a fallback until OpenAI deprecates its base model.

    Calling your model

The comparison uses replies recorded in your traces, so trace before the old model goes away. Serving scales to zero, so the first request after an idle spell can take 2 to 7 minutes. Code on OpenAI's Responses API moves to chat completions first.

What changes in your code?

If your code calls chat completions on the OpenAI SDK, two arguments change. Point the client at Overmind with an Overmind API key, and swap the model id.

Before, calling your OpenAI fine-tune

client = OpenAI()
reply = client.chat.completions.create(model="ft:gpt-4.1-mini-2025-04-14:your-org::abc123", messages=messages)

After, calling the model you trained on Overmind

client = OpenAI(base_url="https://api.overmindlab.ai/api/v1", api_key=os.environ["OVERMIND_API_KEY"])
reply = client.chat.completions.create(model="overmind/<capability-uuid>", messages=messages)

The alias overmind/<capability-uuid> follows whichever model you make live, so later swaps need no code change. To pin one version, use its ft-<job>-<base> id instead. Code built on OpenAI's Responses API moves to chat completions first. Details in the inference docs.

When should you stay on OpenAI?

  • Your fine-tune is not on the 23 October list, it works, and you will not need to retrain it before 6 January 2027. It keeps serving until OpenAI deprecates its base model.
  • Your job needs a proprietary GPT model, DPO, vision fine-tuning or reinforcement fine-tuning. Overmind trains open-weight models with supervised fine-tuning only.
  • A current OpenAI base model with examples in the prompt matches your fine-tune on your own eval set. Then there is nothing to train.

Staying still has a deadline. After 6 January 2027 any retrain has to happen somewhere else, so plan the move before then.

How do you move off OpenAI Evals?

Existing OpenAI evals become read-only on 31 October 2026, and the Evals dashboard and API shut down on 30 November 2026. OpenAI's migration note points to Promptfoo.

In Overmind, an eval run scores one or more variants against a baseline on a frozen dataset version. Build that dataset from traces of the same traffic, so the expected outputs are what your model really returned. There are five evaluator kinds (LLM judge, deterministic, trajectory, statistical and agentic). See eval runs.

Is Overmind an alternative to OpenAI fine-tuning?

Traces

Datasets

Evals

Train

Serve

Overmind

covered:OpenAI SDK traced
covered:full
covered:full
partial:supervised only
covered:full

OpenAI fine-tuning

new jobs end 6 Jan 2027

partial:stored outputs, 30 days
partial:JSONL upload
partial:shuts 30 Nov 2026
partial:no new jobs from 6 Jan 2027
partial:until base deprecated
coveredpartialnot offered

OpenAI fine-tuning as of 5 October 2026. Existing fine-tunes keep serving until their base model is deprecated; no customer can create new jobs from 6 January 2027.

Where should your fine-tune go?

Decision tree

Question 1 of 3, Which model is your fine-tune built on?

Which model is your fine-tune built on?

Four of the five answers here send you somewhere other than Overmind. Staying on OpenAI is a real option until 6 January 2027.

How does Overmind compare with OpenAI fine-tuning?

Overmind and OpenAI fine-tuning compared, 5 October 2026
OvermindOpenAI fine-tuning
New training jobsAvailable, hosted or self-hostedNone for any customer from 6 January 2027; already closed to some organisations
Models38 open-weight bases from 230M to 72B, including Qwen, Llama, Gemma 4 and GPT-OSSProprietary only. SFT and DPO on gpt-4.1, gpt-4.1-mini and gpt-4.1-nano
MethodsSupervised fine-tuning, LoRA or fullSFT and DPO; vision on gpt-4o; reinforcement fine-tuning on o4-mini
Training dataTraces, file uploads, or Langfuse, LangSmith, Braintrust and Galileo connectorsPrepared JSONL, minimum 10 examples. No trace import documented
WeightsDownloadable checkpoint zipNo documented export
ServingOpenAI-compatible chat completions; scales to zeroOpenAI's API, until the base model is deprecated
EvalsEval runs with five evaluator kinds, on datasets from your tracesRead-only from 31 October 2026, shut down 30 November 2026
Open sourceAGPL-3.0 platform, MIT SDK, Docker Compose self-hostNo self-host option documented
PricingFree and Pro plans plus credits; fine-tuned inference on GPU timePer token. Training gpt-4.1 $25 per 1M; tuned gpt-4.1 inference $3 in, $12 out

What does OpenAI fine-tuning do that Overmind does not?

OpenAI fine-tunes proprietary GPT models, including gpt-4.1, and offers DPO, vision fine-tuning on gpt-4o and reinforcement fine-tuning on o4-mini. Overmind trains open-weight models with supervised fine-tuning only. Existing OpenAI fine-tunes also keep serving on OpenAI's API until their base model is deprecated, with no change to your code. If yours works and you will not retrain it before 6 January 2027, staying is a reasonable choice.

What else do teams ask about moving off OpenAI fine-tuning?

Is OpenAI shutting down fine-tuning?

OpenAI is winding down self-serve fine-tuning. Since 7 May 2026 organisations that had never fine-tuned cannot start, and since 2 July neither can those without fine-tuned inference in the past 60 days. From 6 January 2027 no customer can create new jobs. Existing fine-tuned models keep serving until their base model is deprecated (OpenAI deprecations page, checked 5 October 2026).

Which fine-tuned OpenAI models shut down first?

On 23 October 2026 OpenAI removes ft-gpt-3.5-turbo, ft-gpt-4, ft-gpt-4.1-nano-2025-04-14, ft-o4-mini-2025-04-16, ft-babbage-002 and ft-davinci-002, along with several base models. Applications calling them need a tested replacement before that date; OpenAI lists recommended substitutes on its deprecations page.

Can I download the weights of my OpenAI fine-tuned model?

OpenAI documents no way to export fine-tuned weights. Moving means training a new model on the same job, starting from your original training file and your recent production traffic, then checking it against what the old model returned.

Do I have to change my code to move off OpenAI?

If your code calls chat completions, usually only the base URL, API key and model name. Overmind serves trained models on an OpenAI-compatible chat completions endpoint, so the OpenAI SDK keeps working for those calls. Code built on OpenAI's Responses API has to move to chat completions first.

What happens to my OpenAI evals?

Existing evals become read-only on 31 October 2026 and the Evals dashboard and API shut down on 30 November 2026. Overmind's eval runs score variants against a baseline with five evaluator kinds, on datasets built from your production traces.