# Model Training Hack

## Small Model, Big Maths — $1,000 Overmind Challenge

Fine-tune a language model in Overmind to solve GSM8K’s math word problems. Train on 7,473 worked examples, then measure your model’s accuracy on the established 1,319-question test set.

Your score is the percentage of correct answers minus your model’s size in billions of parameters.

**Deadline: Tuesday, October 6, 2026, at 17:00 PDT.**
**Prize: $1,000 for the highest verified adjusted score.**

### How scoring works

```text
Final score = accuracy (%) − model size (B)
```

75% accuracy with a 1.7B model: 75 − 1.7 = 73.3 points. Accuracy is the percentage of correct answers across all 1,319 test questions.

Fractional billions count: a 1.7B model loses 1.7 points. Use the actual total parameter count, including the base model—not just trainable adapter parameters. Quantization does not reduce the count; mixture-of-experts models count all parameters.

Evaluate using EleutherAI’s `gsm8k` task with:
- Strict-match exact-answer accuracy on the complete test set.
- Five-shot prompting, greedy decoding, and one response per question.
- A maximum of 1,024 generated tokens per question.
- Final answers formatted as #### 42.

Rankings use unrounded submitted scores. Lower parameter count breaks ties during verification.

### Example scores

These are illustrative results:

| Model size | Accuracy | Penalty | Final score |
| --- | ---: | ---: | ---: |
| 1B | 71.50% | −1.00 | 70.50 |
| 3B | 72.00% | −3.00 | 69.00 |
| 7B | 76.00% | −7.00 | 69.00 |

The 1B model wins despite having lower raw accuracy.

### What you need to do

1. Choose a model supported by Overmind and fine-tune it. You have until the submission deadline to complete your training and evaluation. Use each training question as the input and its worked answer as the target.
2. Use only the official training split for fine-tuning. Reserve training examples for validation; do not use test answers for training or checkpoint selection, or start from a checkpoint already specifically fine-tuned on GSM8K.
3. Evaluate the model directly, without external calculators, tools, retrieval, ensembles, or repeated sampling.
4. Submit your name, email, and final score before the deadline. Keep your accuracy, total parameter count, starting checkpoint, Overmind training-run reference, final model, and evaluation outputs available for verification.

We will contact the top three teams to verify their scores and parameter counts before confirming the winner.

### Resources

- [Training data — download page](https://github.com/openai/grade-school-math/blob/master/grade_school_math/data/train.jsonl)
- [Test data — download page](https://github.com/openai/grade-school-math/blob/master/grade_school_math/data/test.jsonl)
- [GSM8K dataset and documentation](https://huggingface.co/datasets/openai/gsm8k)
- [Existing GSM8K evaluation task](https://github.com/EleutherAI/lm-evaluation-harness/tree/main/lm_eval/tasks/gsm8k)
- [Evaluation software and setup instructions](https://github.com/EleutherAI/lm-evaluation-harness)

## Submissions

Submit or update your team's score at [the hackathon page](https://www.overmindlab.ai/hack#submissions).

New entries require your name, email address, and final score. Copy and save the private update code shown after submitting to update your score later. Your email is kept private for score verification; no update-code email is sent.

One entry per team. Scores are self-reported and pending verification. Your latest score replaces your previous score. Keep your private update code to edit your entry.

Submissions are open.

Showing 1 of 1 teams. Scores are displayed to 2 decimal places; rankings use unrounded results.

| Rank | Team | Final score |
| --- | --- | ---: |
| 1 | Tyler | 0 |
