Code & Software Engineering
Artificial Analysis public leaderboardSciCode
Scientific programming benchmark results published by Artificial Analysis.
Metric
Accuracy
Results shown
167
Publisher
Artificial Analysis
Snapshot
Fetched 2026-09-08T22:57:37.545Z
167 of 167 models
167 of 167 models
Higher is better
In
Inkling SmallAg
Agnes 2.5 Pro BetaHy
Hy3Qu
Quasar 438B (max)So
Solar Open2 250BIn
InklingAp
Apodex 1.1Ge
Gemma 4 31BRi
Ring-2.6-1TSo
Solar Pro 4St
Step 3.7 FlashNe
Nex-N2-ProAg
Agnes 2.5 Pro AlphaK2
K2 Horizon 375B A23BMo
Motif 3K-
K-EXAONE 2.0Li
Ling 3.0 FlashA.
A.X-K2Tr
Trinity Large ThinkingNe
Nemotron 3 UltraMi
Mistral Medium 3.5No
North Mini CodeMi
Mistral Small 4Co
Command A+Gr
Granite 4.2 30BMe
Mercury 2G9
G9v3-39A5BMi
Mistral Large 3Lo
LongCat 2.0Ne
Nemotron 3 SuperDe
Devstral 2Mi
Mistral Medium 3.1De
Devstral Small 2Ne
Nemotron 3.5 LightningGr
Granite 4.2 8BNe
Nemotron 3 NanoMi
Mistral Small 3.2Mi
Mistral Small 3.1Mi
MiniCPM5-2BSo
Solar Pro 3Gr
Granite 4.2 3BLi
Ling 3.0 TinyMi
Ministral 3 14BGe
Gemma 3 27BCe
Celeris-1Mi
Ministral 3 8BGe
Gemma 3 12BMi
Ministral 3 3BLF
LFM2.5-2.6BList bars visualize relative performance. Plot bars use published 0–100 scores, shrink to a 1.5rem minimum, then scroll horizontally.
Scores are published snapshots and are only comparable within the same benchmark, version, metric, and evaluation configuration.