How scoring works
Final score = accuracy (%) − model size (B)
75% accuracy with a 1.7B model: 75 − 1.7 = 73.3 points.
Accuracy is the percentage of correct answers across all 1,319 test questions.
Fractional billions count: a 1.7B model loses 1.7 points. Use the actual total parameter count, including the base model—not just trainable adapter parameters. Quantization does not reduce the count; mixture-of-experts models count all parameters.
Evaluate using EleutherAI’s gsm8k task with:
- Strict-match exact-answer accuracy on the complete test set.
- Five-shot prompting, greedy decoding, and one response per question.
- A maximum of 1,024 generated tokens per question.
- Final answers formatted as #### 42.
Rankings use unrounded submitted scores. Lower parameter count breaks ties during verification.
Example scores
These are illustrative results:
| Model size | Accuracy | Penalty | Final score |
|---|---|---|---|
| 1B | 71.50% | −1.00 | 70.50 |
| 3B | 72.00% | −3.00 | 69.00 |
| 7B | 76.00% | −7.00 | 69.00 |
The 1B model wins despite having lower raw accuracy.
What you need to do
- Choose a model supported by Overmind and fine-tune it. You have until the submission deadline to complete your training and evaluation. Use each training question as the input and its worked answer as the target.
- Use only the official training split for fine-tuning. Reserve training examples for validation; do not use test answers for training or checkpoint selection, or start from a checkpoint already specifically fine-tuned on GSM8K.
- Evaluate the model directly, without external calculators, tools, retrieval, ensembles, or repeated sampling.
- Submit your name, email, and final score before the deadline. Keep your accuracy, total parameter count, starting checkpoint, Overmind training-run reference, final model, and evaluation outputs available for verification.
We will contact the top three teams to verify their scores and parameter counts before confirming the winner.