# Coding agents

> Canonical: https://www.overmindlab.ai/product/mcp

*Your frontier lab in a box*

Give your coding agent access to your production data, evaluations and model training. Go from a question to a better agent, in one conversation.

Guided example with sample data. Background jobs are condensed; results are illustrative excerpts.

### Analyze: Analyze failures

Inspect traces and failed checks.

**You:** Our dispute agent is contesting cases it should investigate. Look at recent failures against our merchant playbook. What is it missing?

**Your coding agent:** I’ll inspect the failed policy checks and the evidence available when each decision was made.

#### inspect_capability_health

Checked Dispute Resolver · last 14 days

Input:

```json
{
  "capability": "9b1e4f7a-2c3d-4e5f-8a90-1b2c3d4e5f60",
  "days": 14
}
```

Example result (excerpt):

```json
{
  "summary": "Evidence handling is a recurring failure in dispute decisions.",
  "days": 14,
  "capability": "9b1e4f7a-2c3d-4e5f-8a90-1b2c3d4e5f60"
}
```

#### query_failures

Found 12 recent policy failures

Input:

```json
{
  "capability": "9b1e4f7a-2c3d-4e5f-8a90-1b2c3d4e5f60",
  "since_days": 14,
  "limit": 20
}
```

Example result (excerpt):

```json
{
  "n": 12,
  "failures": [
    {
      "trace_id": "a903148c1d2946e197a72184bd096739",
      "summary": "Contested a non-receipt claim despite mismatched delivery evidence.",
      "failed_scores": [
        {
          "evaluator": "Policy adherence",
          "reasoning": "Delivery evidence must match the order address. A delivery status alone is insufficient to contest."
        }
      ]
    }
  ]
}
```

**Your coding agent:** I found 12 recent policy failures. I’m opening a non-receipt case to check whether the evidence was missing—or the model ignored it.

#### query_traces

Opened the decision and its supporting tool calls

Input:

```json
{
  "capability": "9b1e4f7a-2c3d-4e5f-8a90-1b2c3d4e5f60",
  "trace_id": "a903148c1d2946e197a72184bd096739",
  "all_spans": true,
  "limit": 20
}
```

Example result (excerpt):

```json
{
  "summary": "4 spans returned.",
  "n": 4,
  "all_spans": true
}
```

#### resources/read

Read the evidence used in the failed decision

Input:

```json
{
  "uri": "overmind://traces/a903148c1d2946e197a72184bd096739"
}
```

Example result (excerpt):

```json
{
  "kind": "trace",
  "trace_id": "a903148c1d2946e197a72184bd096739",
  "span_count": 4,
  "spans": [
    {
      "name": "resolve_dispute",
      "attributes": {
        "input.value": "A $184 order is marked delivered. The carrier scan is at a different postcode, and there is no signature. The customer denies receiving it.",
        "output.value": {
          "decision": "contest",
          "rationale": "The carrier confirms delivery. Submit the tracking record as evidence."
        }
      }
    },
    {
      "name": "lookup_delivery",
      "attributes": {
        "output.value": {
          "status": "delivered",
          "postcode_match": false,
          "signature": null
        }
      }
    }
  ]
}
```

**The evidence was available; the decision ignored it.** In GLD-042, the agent recommended contesting a non-receipt claim.

- **Carrier status:** Delivered.
- **Delivery location:** Postcode mismatch.
- **Signature:** Not available.

The merchant playbook says:

> Delivery evidence must match the order address. A delivery status alone is insufficient to contest.

The correct next step was to **request matching proof of delivery**.

### Prepare data: Prepare data

Build a reviewed training dataset.

**You:** Use the cases our reviewers have resolved to teach it those distinctions. Keep the Golden cases out of training.

**Your coding agent:** I’ll keep the reviewer-approved resolution as the target, with the original evidence as input. Cases without a reviewed answer will be excluded.

#### create_dataset_from_traces

Collected 4,812 resolved-case traces

Input:

```json
{
  "name": "Dispute Resolutions - Train",
  "capability": "9b1e4f7a-2c3d-4e5f-8a90-1b2c3d4e5f60",
  "filters": {
    "capability": "9b1e4f7a-2c3d-4e5f-8a90-1b2c3d4e5f60"
  },
  "intent": "train",
  "limit": 4812
}
```

Example result (excerpt):

```json
{
  "summary": "Dataset created from 4,812 traces.",
  "traces": 4812,
  "dataset": {
    "id": "c7d1e9a2-5b3f-4e8c-a1d4-6f2b9e0c3a57",
    "name": "Dispute Resolutions - Train",
    "state": "landing"
  },
  "job": {
    "kind": "dataset_run",
    "id": "c7d1e9a2-5b3f-4e8c-a1d4-6f2b9e0c3a57",
    "status": "landing"
  }
}
```

#### get_job

Source data ready

Input:

```json
{
  "kind": "dataset_run",
  "id": "c7d1e9a2-5b3f-4e8c-a1d4-6f2b9e0c3a57"
}
```

Example result (excerpt):

```json
{
  "kind": "dataset_run",
  "id": "c7d1e9a2-5b3f-4e8c-a1d4-6f2b9e0c3a57",
  "label": "Source data ready",
  "status": "completed"
}
```

**Your coding agent:** There are 4,812 source conversations. I’m preserving the reviewed decisions, grouping by case ID and excluding the Golden set before preparation.

#### message_dataset_agent

Prepared evidence → reviewed decision examples

Input:

```json
{
  "dataset": "c7d1e9a2-5b3f-4e8c-a1d4-6f2b9e0c3a57",
  "message": "Drop empty records and deduplicate by case ID. Normalize dates and amounts without dropping original values. Redact personal data. Keep only cases with reviewer-approved resolutions in captured metadata, using that existing resolution as the target. Preserve the original evidence and policy. Exclude IDs and content that overlap Dispute Resolutions - Golden. Do not invent, relabel or replace answers."
}
```

Example result (excerpt):

```json
{
  "summary": "Preparation started with captured reviewer resolutions; no semantic relabelling requested.",
  "job": {
    "kind": "dataset_run",
    "id": "c7d1e9a2-5b3f-4e8c-a1d4-6f2b9e0c3a57",
    "status": "running"
  }
}
```

#### get_job

4,400 examples prepared; Golden cases excluded

Input:

```json
{
  "kind": "dataset_run",
  "id": "c7d1e9a2-5b3f-4e8c-a1d4-6f2b9e0c3a57"
}
```

Example result (excerpt):

```json
{
  "kind": "dataset_run",
  "id": "c7d1e9a2-5b3f-4e8c-a1d4-6f2b9e0c3a57",
  "label": "4,400 examples prepared; Golden cases excluded",
  "status": "completed"
}
```

**Dispute Resolutions — Train:** 4,400 reviewed examples. 250 Golden cases held out; no overlap.

| Split | Cases | Purpose |
| --- | --- | --- |
| Training | 3,520 | Learn the reviewed decisions |
| Validation | 880 | Monitor loss during training |
| Golden | 250 | Compare models after training |

- **TRN-318:** Carrier scan at neighbour’s address; $76 order; no customer confirmation.
  Reviewed target: Request proof of delivery to the correct address before deciding. Source: reviewer-approved resolution, case TRN-318.

- **TRN-562:** $29 renewal on 9 July; cancellation confirmed 4 July; no prior credit.
  Reviewed target: Accept the dispute and recommend one $29 credit. Cancellation preceded the renewal. Source: reviewer-approved resolution, case TRN-562.

- **TRN-904:** Merchant and customer records disagree on the cancellation time; timezone absent.
  Reviewed target: Escalate for source verification. The available records cannot establish event order. Source: reviewer-approved resolution, case TRN-904.

### Train: Train a model

Fine-tune and monitor the run.

**You:** Can a smaller model learn this playbook? Check the data and estimate a training run first.

**Your coding agent:** I’ll check model compatibility and estimate a three-epoch LoRA run. Validation will monitor learning; the Golden evaluation will measure decision quality afterwards.

#### check_finetune_readiness

Training data compatible · zero Golden overlap

Input:

```json
{
  "dataset": "c7d1e9a2-5b3f-4e8c-a1d4-6f2b9e0c3a57",
  "capability": "9b1e4f7a-2c3d-4e5f-8a90-1b2c3d4e5f60"
}
```

Example result (excerpt):

```json
{
  "summary": "Ready for fine-tuning.",
  "ready": true,
  "missing": [],
  "warnings": [],
  "overlap_count": 0,
  "n_candidates": 8,
  "task_type": "structured_output",
  "has_tool_calling": false
}
```

#### estimate_finetune

Qwen3.5 9B · approximately 21 minutes

Input:

```json
{
  "dataset": "c7d1e9a2-5b3f-4e8c-a1d4-6f2b9e0c3a57",
  "base_model": "Qwen/Qwen3.5-9B",
  "n_epochs": 3,
  "use_lora": true
}
```

Example result (excerpt):

```json
{
  "summary": "Estimated training time: 21 minutes.",
  "time_estimate": {
    "seconds": 1260,
    "human": "21 minutes"
  },
  "trained_tokens": 1207560
}
```

**Training recommendation:** Qwen3.5 9B, LoRA, 3 epochs, estimated 21 minutes. Golden cases stay held out.

**You:** Go ahead with that run. Keep the evaluation data held out.

#### start_finetune

Started dispute-resolver-sft-1

Input:

```json
{
  "name": "dispute-resolver-sft-1",
  "dataset": "c7d1e9a2-5b3f-4e8c-a1d4-6f2b9e0c3a57",
  "capability": "9b1e4f7a-2c3d-4e5f-8a90-1b2c3d4e5f60",
  "base_model": "Qwen/Qwen3.5-9B",
  "baseline_model": "gpt-4.1",
  "eval_dataset": "a2f6c8e1-9d4b-4c7a-b3e5-8f1d2c6a9b40",
  "eval_set": "Dispute resolution quality",
  "validation_enabled": true,
  "validation_split_ratio": 0.2,
  "eval_incumbent_before": true,
  "eval_model_before": false,
  "eval_model_after": true,
  "hyperparameters": {
    "n_epochs": 3,
    "use_lora": true,
    "learning_rate": 0.00005,
    "context_length": 4096
  }
}
```

Example result (excerpt):

```json
{
  "summary": "Fine-tuning job started with incumbent-before and trained-after evaluations.",
  "job": {
    "kind": "finetune_job",
    "id": "8c21f7a0-3e5d-4b9f-a2c6-7d1e0f4b8a93",
    "name": "dispute-resolver-sft-1",
    "status": "pending"
  }
}
```

#### get_job

Following preparation and training progress

Input:

```json
{
  "kind": "finetune_job",
  "id": "8c21f7a0-3e5d-4b9f-a2c6-7d1e0f4b8a93"
}
```

Example result (excerpt):

```json
{
  "kind": "finetune_job",
  "id": "8c21f7a0-3e5d-4b9f-a2c6-7d1e0f4b8a93",
  "label": "Following preparation and training progress",
  "status": "running",
  "progress": {
    "phase": "preparing",
    "trained_steps": 0,
    "total_steps": 1200,
    "current_epoch": 0,
    "elapsed_seconds": 0
  }
}
```

**Your coding agent:** The job is queued. I’ll follow preparation, training metrics and validation checkpoints as Overmind reports progress.

**Training run:** dispute-resolver-sft-1. 1,200 optimizer steps, 600 logged metric points, 1,207,560 tokens, 20m 47s. Train loss 0.0031. Sample launch-video metrics; time condensed.

- Epoch 1, 400 steps: validation loss 0.081, token accuracy 93.1%.

- Epoch 2, 800 steps: validation loss 0.034, token accuracy 96.8%.

- Epoch 3, 1200 steps: validation loss 0.0198, token accuracy 97.9%.

#### resources/read

Read metrics history, validation checkpoints and run events

Input:

```json
{
  "uri": "overmind://finetunes/8c21f7a0-3e5d-4b9f-a2c6-7d1e0f4b8a93"
}
```

Example result (excerpt):

```json
{
  "kind": "finetune",
  "id": "8c21f7a0-3e5d-4b9f-a2c6-7d1e0f4b8a93",
  "status": "succeeded",
  "base_model": "Qwen/Qwen3.5-9B",
  "progress": {
    "phase": "completed",
    "trained_steps": 1200,
    "total_steps": 1200,
    "train_loss": 0.0031,
    "token_accuracy": 0.998,
    "eval_token_accuracy": 0.979,
    "tokens_processed": 1207560,
    "elapsed_seconds": 1247,
    "judge_evals": [
      {
        "kind": "baseline",
        "eval_run_id": "0c9b8a7d-6e5f-4a3b-8c2d-1e0f9a8b7c6d",
        "status": "succeeded",
        "sample_count": 250
      },
      {
        "kind": "final",
        "eval_run_id": "7a6b5c4d-3e2f-4a1b-9c8d-7e6f5a4b3c2d",
        "status": "running"
      }
    ],
    "metrics_history": [
      {
        "step": 2,
        "epoch": 0.01,
        "train_loss": 0.9204,
        "token_accuracy": 0.658,
        "lr": 0.00005,
        "grad_norm": 2.9829
      },
      {
        "step": 400,
        "epoch": 1,
        "train_loss": 0.1816,
        "token_accuracy": 0.9427,
        "lr": 0.0000375,
        "grad_norm": 1.6024
      },
      {
        "step": 800,
        "epoch": 2,
        "train_loss": 0.0377,
        "token_accuracy": 0.9913,
        "lr": 0.0000125,
        "grad_norm": 0.5936
      },
      {
        "step": 1200,
        "epoch": 3,
        "train_loss": 0.0031,
        "token_accuracy": 0.998,
        "lr": 0,
        "grad_norm": 0.3026
      }
    ],
    "eval_history": [
      {
        "step": 400,
        "epoch": 1,
        "eval_loss": 0.081
      },
      {
        "step": 800,
        "epoch": 2,
        "eval_loss": 0.034
      },
      {
        "step": 1200,
        "epoch": 3,
        "eval_loss": 0.0198
      }
    ]
  },
  "deployed_model": "d781f29c-3135-41a9-b2f7-76c1df55a004",
  "events": [
    {
      "type": "info",
      "message": "Training complete. Adapter saved and deployment ready."
    }
  ]
}
```

**Your coding agent:** Training completed in 20m 47s. Train loss is 0.0031; validation loss is 0.0198. That measures fit to the examples. The held-out evaluation will tell us whether dispute decisions are better.

### Evaluate: Run evaluations

Compare scores and inspect answers.

**You:** How did it compare with GPT-4.1 on the Golden set? Show me the score changes and any mistakes it still makes.

**Your coding agent:** The training plan includes evaluations before and after training. GPT-4.1 has finished; I’ll follow the trained-model evaluation, compare the same three rubrics, then inspect the decisions.

#### check_evaluation_readiness

Golden dataset and all three evaluators ready

Input:

```json
{
  "dataset": "a2f6c8e1-9d4b-4c7a-b3e5-8f1d2c6a9b40",
  "eval_set": "Dispute resolution quality",
  "mode": "generate"
}
```

Example result (excerpt):

```json
{
  "summary": "Evaluation dataset and evaluator bindings are ready.",
  "ready": true
}
```

#### get_job

Incumbent evaluation complete · 250 cases

Input:

```json
{
  "kind": "eval_run",
  "id": "0c9b8a7d-6e5f-4a3b-8c2d-1e0f9a8b7c6d"
}
```

Example result (excerpt):

```json
{
  "kind": "eval_run",
  "id": "0c9b8a7d-6e5f-4a3b-8c2d-1e0f9a8b7c6d",
  "label": "Incumbent evaluation complete · 250 cases",
  "status": "completed"
}
```

#### get_job

Following the trained-model evaluation · same 250 cases

Input:

```json
{
  "kind": "eval_run",
  "id": "7a6b5c4d-3e2f-4a1b-9c8d-7e6f5a4b3c2d"
}
```

Example result (excerpt):

```json
{
  "kind": "eval_run",
  "id": "7a6b5c4d-3e2f-4a1b-9c8d-7e6f5a4b3c2d",
  "label": "Following the trained-model evaluation · same 250 cases",
  "status": "running"
}
```

**Your coding agent:** The incumbent has finished. I’m waiting for the trained-model run to complete, then I’ll read the comparison and individual answers.

#### get_job

Trained-model evaluation completed

Input:

```json
{
  "kind": "eval_run",
  "id": "7a6b5c4d-3e2f-4a1b-9c8d-7e6f5a4b3c2d"
}
```

Example result (excerpt):

```json
{
  "kind": "eval_run",
  "id": "7a6b5c4d-3e2f-4a1b-9c8d-7e6f5a4b3c2d",
  "label": "Trained-model evaluation completed",
  "status": "completed"
}
```

#### compare_evaluations

Mean rubric score 72.0 → 86.2 · +14.2 points

Input:

```json
{
  "run": "7a6b5c4d-3e2f-4a1b-9c8d-7e6f5a4b3c2d",
  "baseline": "0c9b8a7d-6e5f-4a3b-8c2d-1e0f9a8b7c6d"
}
```

Example result (excerpt):

```json
{
  "summary": "Evaluation comparison ready.",
  "current_run_id": "7a6b5c4d-3e2f-4a1b-9c8d-7e6f5a4b3c2d",
  "baseline_run_id": "0c9b8a7d-6e5f-4a3b-8c2d-1e0f9a8b7c6d",
  "current_name": "Dispute Resolver — trained",
  "baseline_name": "Dispute Resolver — incumbent",
  "rows": [
    {
      "name": "Resolution correctness",
      "current": {
        "mean": 0.89,
        "pass_rate": null,
        "n": 250,
        "variant_count": 1,
        "primary": 0.89
      },
      "baseline": {
        "mean": 0.74,
        "pass_rate": null,
        "n": 250,
        "variant_count": 1,
        "primary": 0.74
      },
      "delta": 0.15,
      "status": "improved"
    },
    {
      "name": "Policy adherence",
      "current": {
        "mean": 0.85,
        "pass_rate": null,
        "n": 250,
        "variant_count": 1,
        "primary": 0.85
      },
      "baseline": {
        "mean": 0.69,
        "pass_rate": null,
        "n": 250,
        "variant_count": 1,
        "primary": 0.69
      },
      "delta": 0.16,
      "status": "improved"
    },
    {
      "name": "Tone",
      "current": {
        "mean": 0.846,
        "pass_rate": null,
        "n": 250,
        "variant_count": 1,
        "primary": 0.846
      },
      "baseline": {
        "mean": 0.73,
        "pass_rate": null,
        "n": 250,
        "variant_count": 1,
        "primary": 0.73
      },
      "delta": 0.116,
      "status": "improved"
    }
  ],
  "overall": {
    "current": {
      "mean": 0.862,
      "pass_rate": null,
      "n": 750,
      "variant_count": 3,
      "primary": 0.862
    },
    "baseline": {
      "mean": 0.7200000000000001,
      "pass_rate": null,
      "n": 750,
      "variant_count": 3,
      "primary": 0.7200000000000001
    },
    "delta": 0.142,
    "status": "improved"
  },
  "trust": {
    "current": {
      "trusted": true,
      "degraded": 0,
      "evaluator_errors": 0,
      "errored": 0,
      "not_applicable": 0
    },
    "baseline": {
      "trusted": true,
      "degraded": 0,
      "evaluator_errors": 0,
      "errored": 0,
      "not_applicable": 0
    }
  },
  "not_applicable_by_evaluator": {}
}
```

#### resources/read

Read incumbent outputs and judge reasoning

Input:

```json
{
  "uri": "overmind://eval-runs/0c9b8a7d-6e5f-4a3b-8c2d-1e0f9a8b7c6d"
}
```

Example result (excerpt):

```json
{
  "kind": "eval_run",
  "id": "0c9b8a7d-6e5f-4a3b-8c2d-1e0f9a8b7c6d",
  "status": "completed",
  "sample_count": 250,
  "samples_truncated": true,
  "samples": [
    {
      "row_index": 0,
      "variant": "GPT-4.1",
      "io": {
        "input": "A $184 order is marked delivered. The carrier scan is at a different postcode, and there is no signature. The customer denies receiving it.",
        "output": "The carrier confirms delivery. Submit the tracking record as evidence."
      },
      "scores": [
        {
          "name": "Resolution correctness",
          "value": 0.25,
          "passed": false,
          "reasoning": "Delivery evidence must match the order address. A delivery status alone is insufficient to contest."
        }
      ]
    },
    {
      "row_index": 1,
      "variant": "GPT-4.1",
      "io": {
        "input": "A $49 renewal was charged on 18 August. Support confirmed cancellation on 12 August. The merchant has not issued a credit.",
        "output": "The customer accepted recurring billing when they subscribed. Contest with the subscription agreement."
      },
      "scores": [
        {
          "name": "Resolution correctness",
          "value": 0.2,
          "passed": false,
          "reasoning": "Accept liability for a renewal after a confirmed cancellation; check for prior credits before recommending a refund."
        }
      ]
    },
    {
      "row_index": 2,
      "variant": "GPT-4.1",
      "io": {
        "input": "A $92 order was delivered to the order address. A signed receipt and the customer’s message confirming delivery are attached.",
        "output": "Contest with the signed delivery receipt and the customer’s confirmation."
      },
      "scores": [
        {
          "name": "Resolution correctness",
          "value": 0.95,
          "passed": true,
          "reasoning": "Contest a non-receipt claim when matching delivery evidence and customer confirmation are available."
        }
      ]
    },
    {
      "row_index": 3,
      "variant": "GPT-4.1",
      "io": {
        "input": "The cancellation email and billing event have the same timestamp but different timezones. Their ordering cannot be established from the captured records.",
        "output": "The charge appears to predate cancellation. Contest it."
      },
      "scores": [
        {
          "name": "Resolution correctness",
          "value": 0.2,
          "passed": false,
          "reasoning": "Escalate unresolved conflicts in source evidence. Do not infer the event ordering."
        }
      ]
    }
  ]
}
```

#### resources/read

Read improved cases and the remaining evidence-ordering failure

Input:

```json
{
  "uri": "overmind://eval-runs/7a6b5c4d-3e2f-4a1b-9c8d-7e6f5a4b3c2d"
}
```

Example result (excerpt):

```json
{
  "kind": "eval_run",
  "id": "7a6b5c4d-3e2f-4a1b-9c8d-7e6f5a4b3c2d",
  "status": "completed",
  "sample_count": 250,
  "samples_truncated": true,
  "samples": [
    {
      "row_index": 0,
      "variant": "Dispute Resolver SFT",
      "io": {
        "input": "A $184 order is marked delivered. The carrier scan is at a different postcode, and there is no signature. The customer denies receiving it.",
        "output": "The scan does not match the order address. Request proof of delivery before deciding whether to contest."
      },
      "scores": [
        {
          "name": "Resolution correctness",
          "value": 1,
          "passed": true,
          "reasoning": "The trained response uses the postcode mismatch and missing signature. The original response treats a status label as sufficient evidence."
        }
      ]
    },
    {
      "row_index": 1,
      "variant": "Dispute Resolver SFT",
      "io": {
        "input": "A $49 renewal was charged on 18 August. Support confirmed cancellation on 12 August. The merchant has not issued a credit.",
        "output": "Cancellation was confirmed six days before renewal. Accept liability and recommend a $49 refund; no earlier credit exists."
      },
      "scores": [
        {
          "name": "Resolution correctness",
          "value": 0.95,
          "passed": true,
          "reasoning": "The cancellation timestamp takes precedence over the original billing agreement. The trained response also checks for a duplicate credit."
        }
      ]
    },
    {
      "row_index": 2,
      "variant": "Dispute Resolver SFT",
      "io": {
        "input": "A $92 order was delivered to the order address. A signed receipt and the customer’s message confirming delivery are attached.",
        "output": "The evidence is complete and consistent. Prepare a response citing the signed receipt and the delivery confirmation."
      },
      "scores": [
        {
          "name": "Resolution correctness",
          "value": 0.95,
          "passed": true,
          "reasoning": "Both responses select the correct resolution and cite the available evidence. This is a control case, not a new improvement."
        }
      ]
    },
    {
      "row_index": 3,
      "variant": "Dispute Resolver SFT",
      "io": {
        "input": "The cancellation email and billing event have the same timestamp but different timezones. Their ordering cannot be established from the captured records.",
        "output": "The cancellation appears to have arrived first. Accept the dispute."
      },
      "scores": [
        {
          "name": "Resolution correctness",
          "value": 0.35,
          "passed": false,
          "reasoning": "Both models infer an order the evidence does not establish. Keep this case in the next review batch; the trained model is not correct on every case."
        }
      ]
    }
  ]
}
```

**Paired evaluation outputs:** same held-out cases and policy checks.

Mean rubric score: 72.0 → 86.2 / 100 (+14.2 points). These are illustrative scores, not pass rates.

| Metric | GPT-4.1 | Trained | Change |
| --- | --- | --- | --- |
| Resolution correctness | 74.0 | 89.0 | +15.0 |
| Policy adherence | 69.0 | 85.0 | +16.0 |
| Tone | 73.0 | 84.6 | +11.6 |
| Mean score | 72.0 | 86.2 | +14.2 |

- Delivery exception: current model — Contest; trained — Request evidence; reference — Request matching proof of delivery; do not contest yet.
  Judge reasoning: The trained response uses the postcode mismatch and missing signature. The original response treats a status label as sufficient evidence.

- Cancelled subscription: current model — Contest; trained — Accept liability; reference — Accept liability and recommend a single $49 refund.
  Judge reasoning: The cancellation timestamp takes precedence over the original billing agreement. The trained response also checks for a duplicate credit.

- Evidence complete: current model — Contest; trained — Contest; reference — Contest with both supporting documents.
  Judge reasoning: Both responses select the correct resolution and cite the available evidence. This is a control case, not a new improvement.

- Still needs review: current model — Contest; trained — Accept liability; reference — Escalate for review of the original timestamps.
  Judge reasoning: Both models infer an order the evidence does not establish. Keep this case in the next review batch; the trained model is not correct on every case.

### Run inference: Run inference

Test the model and update the code.

**You:** Try a new delivery dispute. If it follows the playbook, activate it for Dispute Resolver and show me the code change. Keep human approval for financial actions.

**Your coding agent:** I’ll test a new case against the ready deployment first. The model will return a recommendation; the existing approval step remains in your agent.

#### get_job

Deployment ready for inference

Input:

```json
{
  "kind": "deployment",
  "id": "d781f29c-3135-41a9-b2f7-76c1df55a004"
}
```

Example result (excerpt):

```json
{
  "kind": "deployment",
  "id": "d781f29c-3135-41a9-b2f7-76c1df55a004",
  "label": "Deployment ready for inference",
  "status": "ready"
}
```

#### run_inference

New case → request matching delivery evidence

Input:

```json
{
  "deployment": "d781f29c-3135-41a9-b2f7-76c1df55a004",
  "messages": [
    {
      "role": "user",
      "content": "Triage case DSP-1284 using the merchant playbook. A $136 order is marked delivered, but the delivery photo shows a different door number. No signature is available. The customer denies receipt. Return a JSON recommendation; do not execute a financial action."
    }
  ],
  "temperature": 0,
  "max_tokens": 512
}
```

Example result (excerpt):

```json
{
  "summary": "Inference completed.",
  "model_id": "overmind/8c21f7a0-3e5d-4b9f-a2c6-7d1e0f4b8a93",
  "content": "{\"case_id\":\"DSP-1284\",\"decision\":\"request_evidence\",\"reason\":\"The delivery photo does not match the order address, so delivery is not established.\",\"evidence_used\":[\"delivery_photo_address_mismatch\",\"signature_missing\"],\"next_action\":\"Request proof of delivery to the order address before deciding whether to contest.\",\"requires_human_approval\":true}",
  "usage": {
    "prompt_tokens": 218,
    "completion_tokens": 96,
    "total_tokens": 314
  },
  "latency_ms": 1840,
  "is_cold": false,
  "finish_reason": "stop",
  "truncated": false
}
```

**New inference request:** Triage case DSP-1284 using the merchant playbook. A $136 order is marked delivered, but the delivery photo shows a different door number. No signature is available. The customer denies receipt. Return a JSON recommendation; do not execute a financial action.

**Trained model response:**

```json
{
  "case_id": "DSP-1284",
  "decision": "request_evidence",
  "reason": "The delivery photo does not match the order address, so delivery is not established.",
  "evidence_used": [
    "delivery_photo_address_mismatch",
    "signature_missing"
  ],
  "next_action": "Request proof of delivery to the order address before deciding whether to contest.",
  "requires_human_approval": true
}
```

**Your coding agent:** It requests the missing proof of delivery instead of contesting automatically. I’ll activate this model and prepare the alias change.

#### set_active_model

Activated the trained Dispute Resolver

Input:

```json
{
  "capability": "9b1e4f7a-2c3d-4e5f-8a90-1b2c3d4e5f60",
  "deployment": "d781f29c-3135-41a9-b2f7-76c1df55a004"
}
```

Example result (excerpt):

```json
{
  "summary": "Active model updated for Dispute Resolver.",
  "cleared": false,
  "active_model": {
    "id": "d781f29c-3135-41a9-b2f7-76c1df55a004",
    "model_id": "overmind/8c21f7a0-3e5d-4b9f-a2c6-7d1e0f4b8a93",
    "status": "ready"
  }
}
```

#### get_model_swap_prompt

Prepared the permanent capability alias

Input:

```json
{
  "finetune": "8c21f7a0-3e5d-4b9f-a2c6-7d1e0f4b8a93",
  "pin": false
}
```

Example result (excerpt):

```json
{
  "summary": "Model-swap instructions prepared.",
  "capability_id": "9b1e4f7a-2c3d-4e5f-8a90-1b2c3d4e5f60",
  "old_model": "gpt-4.1",
  "new_model": "overmind/9b1e4f7a-2c3d-4e5f-8a90-1b2c3d4e5f60",
  "pin": false
}
```

**Proposed integration:** agents/dispute_resolver.py

```python
from openai import OpenAI
import os

client = OpenAI(
    base_url="https://api.overmindlab.ai/api/v1",
    api_key=os.environ["OVERMIND_API_KEY"],
)

response = client.chat.completions.create(
    model="overmind/9b1e4f7a-2c3d-4e5f-8a90-1b2c3d4e5f60",
    messages=messages,
)
```

**Your coding agent:** The new case follows the playbook. The trained model is active, and the code change is ready for review. Future model swaps can use this same capability alias.

## Connect your coding agent.

Choose your coding agent and copy the prompt to connect your project to Overmind Cloud.

[Open Console](https://console.overmindlab.ai) · [Read the quickstart](https://docs.overmindlab.ai/quickstart)

### Onboarding prompt for Cursor

```text
Connect Cursor to Overmind Cloud using the hosted MCP server at https://api.overmindlab.ai/api/mcp/.

Use the project already open in this workspace. If no project is open, ask me to open the project I want to connect.

Read https://docs.overmindlab.ai/quickstart. Guide me through signing in at https://console.overmindlab.ai and obtaining the cloud onboarding credentials. Keep the temporary bootstrap key in the current shell session only, out of tracked files and chat output.

Install the published overmind package from PyPI using this project's package manager and dependency manifest. For a non-Python project, use an isolated CLI environment. Set OVERMIND_API_URL=https://api.overmindlab.ai and run overmind init --ide cursor --env production through that environment.

Read references/onboard.md and references/onboarding-progress.md in the installed Overmind skill. Show the onboarding roadmap and what will be synced, then run overmind sync to connect the Console project and configure MCP with its project-scoped credential.

Ask me to reload Cursor once. After reload, read overmind://project/current to verify the connection and continue the installed onboarding skill without repeating bootstrap setup. Summarize the setup and recommend the next step.
```

MCP configuration: `.cursor/mcp.json`

### Onboarding prompt for Claude Code

```text
Connect Claude Code to Overmind Cloud using the hosted MCP server at https://api.overmindlab.ai/api/mcp/.

Use the project already open in this workspace. If no project is open, ask me to open the project I want to connect.

Read https://docs.overmindlab.ai/quickstart. Guide me through signing in at https://console.overmindlab.ai and obtaining the cloud onboarding credentials. Keep the temporary bootstrap key in the current shell session only, out of tracked files and chat output.

Install the published overmind package from PyPI using this project's package manager and dependency manifest. For a non-Python project, use an isolated CLI environment. Set OVERMIND_API_URL=https://api.overmindlab.ai and run overmind init --ide claude --env production through that environment.

Read references/onboard.md and references/onboarding-progress.md in the installed Overmind skill. Show the onboarding roadmap and what will be synced, then run overmind sync to connect the Console project and configure MCP with its project-scoped credential.

Ask me to reload Claude Code once. After reload, read overmind://project/current to verify the connection and continue the installed onboarding skill without repeating bootstrap setup. Summarize the setup and recommend the next step.
```

MCP configuration: `.mcp.json`

### Onboarding prompt for Codex

```text
Connect Codex to Overmind Cloud using the hosted MCP server at https://api.overmindlab.ai/api/mcp/.

Use the project already open in this workspace. If no project is open, ask me to open the project I want to connect.

Read https://docs.overmindlab.ai/quickstart. Guide me through signing in at https://console.overmindlab.ai and obtaining the cloud onboarding credentials. Keep the temporary bootstrap key in the current shell session only, out of tracked files and chat output.

Install the published overmind package from PyPI using this project's package manager and dependency manifest. For a non-Python project, use an isolated CLI environment. Set OVERMIND_API_URL=https://api.overmindlab.ai and run overmind init --ide codex --env production through that environment.

Read references/onboard.md and references/onboarding-progress.md in the installed Overmind skill. Show the onboarding roadmap and what will be synced, then run overmind sync to connect the Console project and configure MCP with its project-scoped credential.

Ask me to reload Codex once. After reload, read overmind://project/current to verify the connection and continue the installed onboarding skill without repeating bootstrap setup. Summarize the setup and recommend the next step.
```

MCP configuration: `.codex/config.toml`

### Onboarding prompt for OpenCode

```text
Connect OpenCode to Overmind Cloud using the hosted MCP server at https://api.overmindlab.ai/api/mcp/.

Use the project already open in this workspace. If no project is open, ask me to open the project I want to connect.

Read https://docs.overmindlab.ai/quickstart. Guide me through signing in at https://console.overmindlab.ai and obtaining the cloud onboarding credentials. Keep the temporary bootstrap key in the current shell session only, out of tracked files and chat output.

Install the published overmind package from PyPI using this project's package manager and dependency manifest. For a non-Python project, use an isolated CLI environment. Set OVERMIND_API_URL=https://api.overmindlab.ai and run overmind init --ide opencode --env production through that environment.

Read references/onboard.md and references/onboarding-progress.md in the installed Overmind skill. Show the onboarding roadmap and what will be synced, then run overmind sync to connect the Console project and configure MCP with its project-scoped credential.

Ask me to reload OpenCode once. After reload, read overmind://project/current to verify the connection and continue the installed onboarding skill without repeating bootstrap setup. Summarize the setup and recommend the next step.
```

MCP configuration: `opencode.json`