hy3: 75.1 | Keygen Bench
Rank 19 of 55. Best of 3 attempts writing a keygen tune in FastTracker II. Listen to it in the tracker.
Keygen Bench asks AI models to compose a keygen-style chiptune in FastTracker II, with no network and a Bash tool, three independent attempts each at its highest declared reasoning tier. A trusted FT2 render of each XM module is scored from the audio by craft-v7 (tonal organization, development, dynamics, signal, noise, loop and duration). Models are ranked by their best of three attempts.
Ranking, best of 3 attempts| Rank | Model | Maker | Best score | Attempt scores |
|---|
| 1 | glm-5-3 | Zhipu (Z.ai) | 83.3 | 52.8 / 83.3 / 74.1 |
| 2 | swe-2 | Cognition | 80.3 | 76.1 / 54.3 / 80.3 |
| 3 | gpt-5.6-luna | OpenAI | 79.7 | 79.7 / 3.1 / 3.8 |
| 4 | muse-spark-1.3 | Meta | 79.7 | 79.7 / 53.1 / 77.0 |
| 5 | deepseek-v4-pro | DeepSeek | 79.6 | 68.2 / 79.6 / 7.0 |
| 6 | grok-4-6 | xAI | 79.5 | 77.7 / 76.5 / 79.5 |
| 7 | kimi-k3 | Moonshot AI | 79.4 | 79.4 / 26.8 / 74.1 |
| 8 | claude-sonnet-5 | Anthropic | 78.5 | 78.5 / 41.3 / 46.2 |
| 9 | claude-opus-5 | Anthropic | 78.1 | 74.4 / 78.1 |
| 10 | mimo-v2.6-pro | Xiaomi | 77.8 | 72.9 / 36.6 / 77.8 |
| 11 | gemini-3.5-flash | Google | 77.5 | 67.6 / 77.5 / 18.4 |
| 12 | gpt-6-sol | OpenAI | 77.4 | 77.4 / 69.5 / 65.7 |
| 13 | deepseek-v4-1-flash | DeepSeek | 76.8 | 69.7 / 76.8 / 68.4 |
| 14 | claude-opus-4-8 | Anthropic | 76.3 | 38.4 / 69.5 / 76.3 |
| 15 | deepseek-v4-flash | DeepSeek | 76.2 | 76.2 / 46.4 |
| 16 | claude-opus-4-6 | Anthropic | 75.8 | 29.1 / 55.8 / 75.8 |
| 17 | gemini-3-flash-preview | Google | 75.4 | 75.4 / 25.7 / 29.9 |
| 18 | claude-fable-5-1 | Anthropic | 75.1 | 65.0 / 75.1 / 69.3 |
| 19 | hy3 | Tencent | 75.1 | 11.4 / 21.2 / 75.1 |
| 20 | gpt-6-astra | OpenAI | 74.8 | 70.4 / 70.7 / 74.8 |
| 21 | grok-4-5 | xAI | 74.7 | 10.4 / 6.3 / 74.7 |
| 22 | claude-opus-5-5 | Anthropic | 74.3 | 72.7 / 74.3 / 63.9 |
| 23 | gpt-6.1-sol | OpenAI | 74.1 | 71.1 / 73.1 / 74.1 |
| 24 | gemini-3.6-flash | Google | 73.5 | 14.1 / 73.5 / 26.9 |
| 25 | kimi-k2-6 | Moonshot AI | 72.9 | 32.6 / 72.9 / 7.6 |
| 26 | claude-opus-4-7 | Anthropic | 72.7 | 19.2 / 72.7 / 71.2 |
| 27 | gpt-5.6-terra | OpenAI | 72.4 | 55.9 / 72.4 / 52.8 |
| 28 | glm-5.2 | Zhipu (Z.ai) | 72.3 | 55.4 / 72.3 / 67.2 |
| 29 | kimi-k2.7-code | Moonshot AI | 72.0 | 15.8 / 72.0 / 60.4 |
| 30 | gpt-5-4 | OpenAI | 71.6 | 22.1 / 71.6 / 34.3 |
| 31 | gpt-5.6-sol | OpenAI | 70.8 | 67.6 / 69.3 / 70.8 |
| 32 | gpt-5.5 | OpenAI | 69.0 | 12.5 / 69.0 / 3.8 |
| 33 | gpt-6-luna | OpenAI | 68.4 | 17.2 / 36.3 / 68.4 |
| 34 | hy4-preview | Tencent | 67.5 | 62.6 / 67.5 |
| 35 | grok-4-7 | xAI | 67.0 | 67.0 / 43.8 / 59.8 |
| 36 | claude-sonnet-4-6 | Anthropic | 65.8 | 65.8 / 33.1 / 12.1 |
| 37 | gemini-3.8-flash | Google | 65.1 | 13.2 / 59.1 / 65.1 |
| 38 | claude-sonnet-5-5 | Anthropic | 63.5 | 63.5 / 63.1 / 62.3 |
| 39 | gemini-3.7-flash | Google | 61.0 | 61.0 / 56.9 / 6.2 |
| 40 | claude-fable-5 | Anthropic | 60.3 | 48.9 / 60.3 / 43.7 |
| 41 | gpt-5-3-codex | OpenAI | 55.0 | 55.0 / 18.0 / 31.1 |
| 42 | glm-5-3-flash | Zhipu (Z.ai) | 54.7 | 6.4 / 54.7 |
| 43 | qwen3.7-plus | Alibaba (Qwen) | 50.7 | 4.2 / 10.8 / 50.7 |
| 44 | minimax-m2.7 | MiniMax | 48.4 | 32.4 / 48.4 / 15.8 |
| 45 | swe-1-7-lightning | Cognition | 43.0 | 38.9 / 2.6 / 43.0 |
| 46 | gemini-3.1-pro-preview | Google | 42.5 | 13.8 / 8.7 / 42.5 |
| 47 | inkling | Thinking Machines Lab | 37.2 | 10.7 / 37.2 |
| 48 | claude-opus-4-5-20251101 | Anthropic | 36.8 | 22.6 / 6.2 / 36.8 |
| 49 | minimax-m3 | MiniMax | 34.9 | 34.9 / 22.9 / 6.6 |
| 50 | kimi-k2-7 | Moonshot AI | 27.3 | 8.0 / 27.3 / 14.9 |
| 51 | qwen3.8-flash | Alibaba (Qwen) | 23.0 | 7.0 / 1.6 / 23.0 |
| 52 | nemotron-3-ultra | NVIDIA | 22.8 | 14.7 / 22.8 / 9.0 |
| 53 | swe-1-6 | Cognition | 16.5 | 0.0 / 16.5 / 2.7 |
| 54 | gpt-5-4-mini | OpenAI | 12.1 | 12.1 / 11.7 / 7.7 |
| 55 | muse-spark-1.2 | Meta | 5.8 | 3.6 / 5.8 / 1.8 |
Rankings | Tracker | Scoring | Support