Power Ranking
We gave 124 real tickets from a 1.1M loc Rust codebase ↗ to each model and judged the solutions with Astra and Fable.
Intelligence
On each task, Astra/high and Fable/high score every model's submission. Each pair of scores on the same task counts as a win, a loss or a tie; a model that did not finish loses to every model that did. A Bradley-Terry model with ties turns these results into one strength number per model (log scale, 0 is the average). A gap of 1.0 means the stronger model is expected to win about 73 percent of the comparisons that are not ties.
Speed
Working minutes per task. We fit each model's time and each task's size together, so a model is not penalized for getting harder tasks, and we report the time for a task of typical size.
Value
How many tasks $200 USD buys, using token counts adjusted for task size in the same way. API: at the vendor's list prices. Subscription: we read how much of a plan's weekly allowance our runs used, convert that to a weekly budget, and scale it to $200 USD of plan price.
Power
A weighted geometric mean of intelligence, speed and value, each as a fraction of the best model's (see the Power chart, where you can change the weights). Our weights are intelligence 1, speed 0.5 and value 0.5, so intelligence counts as much as speed and value together. A geometric mean multiplies the scores, so a model cannot make up for a very low score on one measure by being a little better on another.
Notes
- Subscription value is normalized to a $200 USD monthly subscription price.
- Hollow points and bars have no subscription figure and use API list prices.
- Bars show 95% confidence intervals.
- API only (no subscription sold): DeepSeek Flash max, DeepSeek Flash high, DeepSeek Flash low. On the subscription setting they are shown at API list prices.
- Muse Spark 1.3 ultra: its subscription budget is unknown (Muse Code Power (20x Everyday), $50 USD a month). On the subscription setting it is shown at API list prices.