# Sources

Every original source behind the DeepSWE 1.1 tracker, numbered as on the page. 40 sources; generated by `npm run build` from `data/deepswe-1.1.csv` and `data/meta.json`.

1. [Datacurve DeepSWE 1.1 leaderboard](https://deepswe.datacurve.ai/) — Datacurve · Official leaderboard · published 15 Jun 2026 · read 9 Oct 2026. Pass@1 for GPT-6 Astra 67%–74.1% across 5 settings; Gemini 3.8 Flash 71%–73.8% across 2 settings; Claude Opus 5 58.1%–73.7% across 5 settings; GPT-5.6 Sol 45.4%–72.7% across 5 settings; Claude Fable 5 59.6%–69.9% across 5 settings; GPT-5.6 Terra 24.1%–69.6% across 5 settings; and 22 more models, with cost per task, tokens, confidence intervals. (70 readings)
2. [Datacurve DeepSWE 1.1 leaderboard, launch snapshot (Wayback)](https://web.archive.org/web/20260615111157/https://deepswe.datacurve.ai/) — Internet Archive (capture of the Datacurve board) · Official leaderboard · published 15 Jun 2026 · read 9 Oct 2026. Pass@1 for Claude Opus 4.7 31.6%–54.2% across 4 settings; Gemini 3.5 Flash 28.3% [medium]; Claude Opus 4.6 27.6% [max]; GPT-5.4 mini 24.3% [xhigh]; Kimi K2.6 23.9%; MiniMax M3 20.4%; and 10 more models, with cost per task, tokens, confidence intervals. (19 readings)
3. [DeepSWE changelog](https://deepswe.datacurve.ai/changelog) — Datacurve · Official changelog · published 3 Sept 2026 · read 9 Oct 2026. Lists each model addition by date; the most recent is GPT-6 Astra (all efforts) on 3 Sep 2026.
4. [Run DeepSWE](https://deepswe.datacurve.ai/run) — Datacurve · Official documentation · published 15 Jun 2026 · read 9 Oct 2026. Official scores are produced with Pier running mini-swe-agent on Modal, in isolated containers.
5. [leaderboard-live.json (v1.1 data file)](https://deepswe.datacurve.ai/artifacts/v1.1/leaderboard-live.json) — Datacurve · Official data file · published 22 Sept 2026 · read 9 Oct 2026. The board's data file: generated 22 Sep 2026, latest job finished 1 Sep; its costs differ from the page's for 19 of 70 configurations.
6. [GLM-5.2 HuggingFace Model Card](https://huggingface.co/zai-org/GLM-5.2) — Z.ai · Lab self-report · published 16 Jun 2026 · read 9 Oct 2026. Pass@1 for GLM-5.2 46.2% [max]. (1 reading)
7. [Poolside Laguna S 2.1 launch post](https://poolside.ai/blog/introducing-laguna-s-2-1) — Poolside · Lab self-report · published 21 Jul 2026 · read 9 Oct 2026. Pass@1 for Laguna S 2.1 40.4% [max]. (1 reading)
8. [Kimi-K3 HuggingFace Model Card](https://huggingface.co/moonshotai/Kimi-K3) — Moonshot AI · Lab self-report · published 23 Jul 2026 · read 9 Oct 2026. Pass@1 for Kimi K3 67.5% [max]. (1 reading)
9. [Claude Opus 5 System Card](https://www-cdn.anthropic.com/c5fbac3f0b1280a933ebd26d3cb8bb9f5bdeaf48/Claude%20Opus%205%20System%20Card.pdf) — Anthropic · Lab self-report · published 24 Jul 2026 · read 9 Oct 2026. Pass@1 for Claude Opus 5 68.8% [max]. (1 reading)
10. [Qwen3.8-27B HuggingFace Model Card](https://huggingface.co/Qwen/Qwen3.8-27B) — Alibaba (Qwen) · Lab self-report · published 5 Aug 2026 · read 9 Oct 2026. Pass@1 for Qwen3.8-27B 42.2%; Qwen3.6-27B 13.3%. (2 readings)
11. [Qwen3.8-2.4T-A95B (Qwen3.8-Max) HuggingFace Model Card](https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B) — Alibaba (Qwen) · Lab self-report · published 8 Aug 2026 · read 9 Oct 2026. Pass@1 for Qwen3.8 Max 56.6%; Qwen3.7 Max 21.6%. (2 readings)
12. [xAI Grok 4.6 Model Card](https://media.x.ai/v1/website/card-4p6-4cd2dc57.pdf) — xAI · Lab self-report · published 18 Aug 2026 · read 9 Oct 2026. Pass@1 for Grok 4.6 65.9% [high]. (1 reading)
13. [Qwen3.8-Flash-Next HuggingFace Model Card](https://huggingface.co/Qwen/Qwen3.8-Flash-Next) — Alibaba (Qwen) · Lab self-report · published 24 Aug 2026 · read 9 Oct 2026. Pass@1 for Qwen3.8-Flash-Next 58.7%; Qwen3.7 Plus 16.5%. (2 readings)
14. [GLM-5.3 HuggingFace Model Card](https://huggingface.co/zai-org/GLM-5.3) — Z.ai · Lab self-report · published 25 Aug 2026 · read 9 Oct 2026. Pass@1 for GLM-5.3 66.9% [max]. (1 reading)
15. [Tencent Hy4-preview Technical Report & Model Card](https://huggingface.co/tencent/Hy4-preview) — Tencent · Lab self-report · published 27 Aug 2026 · read 9 Oct 2026. Pass@1 for Hy4 Preview 64.3%; Hy3 28%. (2 readings)
16. [Claude Fable 5.1 & Claude Mythos 5.1 System Card](https://www-cdn.anthropic.com/0339e6a7c5c7b87f5c07798616dc32c215d14235/Claude%20Fable%205.1%20&%20Claude%20Mythos%205.1%20System%20Card.pdf) — Anthropic · Lab self-report · published 1 Sept 2026 · read 9 Oct 2026. Pass@1 for Claude Fable 5.1 67.4% [max]. (1 reading)
17. [Meta AI Research Muse Spark 1.3 Evaluation Methodology](https://research.meta.ai/static/muse-spark-1-3-multimodal-evaluation-methodology) — Meta · Lab self-report · published 2 Sept 2026 · read 9 Oct 2026. Pass@1 for Muse Spark 1.3 75.4% [max]. (1 reading)
18. [Nex-AGI Nex-N2.5 GitHub README / Model Card](https://github.com/nex-agi/Nex-N2.5) — Nex-AGI · Lab self-report · published 8 Sept 2026 · read 9 Oct 2026. Pass@1 for Nex-N2.5-Max 65.6%; Nex-N2.5-Pro 55.8%; Nex-N2.5-Mini 36.1%. (3 readings)
19. [Cognition SWE-2 launch post](https://cognition.com/blog/swe-2) — Cognition · Lab self-report · published 10 Sept 2026 · read 9 Oct 2026. Pass@1 for SWE-2 73%; SWE-1.7 37.7%. (2 readings)
20. [DeepSeek-V4.1-Flash HuggingFace Model Card](https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash) — DeepSeek · Lab self-report · published 10 Sept 2026 · read 9 Oct 2026. Pass@1 for DeepSeek V4.1 Flash 65.5%–74.2% across 8 settings; DeepSeek V4 Pro 62.7% [max]; DeepSeek V4 Flash 54.4% [max]. (10 readings)
21. [Google Gemini 3.8 Flash launch blog and Model Evaluation Report](https://blog.google/innovation-and-ai/models-and-research/gemini-models/3-8-flash-and-3-8-flash-cyber/) — Google DeepMind · Lab self-report · published 17 Sept 2026 · read 9 Oct 2026. Pass@1 for Gemini 3.8 Flash 73.7% [high], with cost per task, tokens. (1 reading)
22. [StepFun Step 5 Preview Announcement](https://www.stepfun.com/) — StepFun · Lab self-report · published 20 Sept 2026 · read 9 Oct 2026. Pass@1 for Step 5 Preview 67.7% [high]. (1 reading)
23. [xAI Grok 4.7 Model Card](https://media.x.ai/v1/website/card4p7-3a96f40b.pdf) — xAI · Lab self-report · published 21 Sept 2026 · read 9 Oct 2026. Pass@1 for Grok 4.7 71% [high]. (1 reading)
24. [Xiaomi MiMo-V2.6 Launch Announcement and Appendix](https://mimo.xiaomi.com/mimo-v2-6) — Xiaomi · Lab self-report · published 21 Sept 2026 · read 9 Oct 2026. Pass@1 for MiMo-V2.6-Pro 71.9% [max]; MiMo-V2.6-Flash 67.9% [max]. (2 readings)
25. [Claude Opus 5.5 System Card](https://www-cdn.anthropic.com/fc1b44717c85dc068bc6ba5024219938094694bd/Claude%20Opus%205.5%20System%20Card.pdf) — Anthropic · Lab self-report · published 22 Sept 2026 · read 9 Oct 2026. Pass@1 for Claude Opus 5.5 74.2% [max]. (1 reading)
26. [OpenAI GPT-6 Sol and Luna launch post](https://openai.com/index/introducing-gpt-6-sol-and-luna/) — OpenAI · Lab self-report · published 22 Sept 2026 · read 9 Oct 2026. Pass@1 for GPT-6 Sol 37.2%–68.8% across 5 settings; GPT-6 Luna 2.4%–66.6% across 5 settings; GPT-5.6 Luna 1.2%–62.2% across 5 settings, with cost per task. (15 readings)
27. [Fireworks AI Ember-1 Announcement](https://fireworks.ai/blog/ember-1) — Fireworks AI · Lab self-report · published 23 Sept 2026 · read 9 Oct 2026. Pass@1 for Ember-1 75.2% [thinking]; Kimi K3 55.8%–66.4% across 3 settings, with cost per task. (4 readings)
28. [Gemini 4 Argon Model Evaluation Report](https://storage.googleapis.com/deepmind-media/gemini/gemini_4_argon_model_evaluation.pdf) — Google DeepMind · Lab self-report · published 24 Sept 2026 · read 9 Oct 2026. Pass@1 for Gemini 4 Argon 77.9%. (1 reading)
29. [Claude Sonnet 5.5 System Card](https://www-cdn.anthropic.com/870c8f525702625d2c62fc6dd04c857e3250bec1/Claude%20Sonnet%205.5%20System%20Card.pdf) — Anthropic · Lab self-report · published 28 Sept 2026 · read 9 Oct 2026. Pass@1 for Claude Sonnet 5.5 71% [max]. (1 reading)
30. [OpenAI GPT-6.1 Sol launch post](https://openai.com/index/introducing-gpt-6-1-sol/) — OpenAI · Lab self-report · published 29 Sept 2026 · read 9 Oct 2026. Pass@1 for GPT-6.1 Sol 64.4%–75.2% across 5 settings, with cost per task. (5 readings)
31. [Reflection AI Beam launch post](https://reflection.ai/blog/introducing-beam) — Reflection AI · Lab self-report · published 5 Oct 2026 · read 9 Oct 2026. Pass@1 for Beam 44.4%. (1 reading)
32. [Mistral Large 4 Announcement](https://mistral.ai/news/mistral-large-4/) — Mistral AI · Lab self-report · published 6 Oct 2026 · read 9 Oct 2026. Pass@1 for Mistral Large 4 61.7% [thinking]. (1 reading)
33. [Claude Haiku 5.5 System Card](https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf) — Anthropic · Lab system card · published 29 Sept 2026 · read 9 Oct 2026. Reports SWE-bench Pro, FrontierSWE v2 and ProgramBench for Haiku 5.5, and no DeepSWE 1.1 result.
34. [Independent DeepSWE Audit by entrpi](https://entrpi.github.io/misc/deep-swe-minimax-m3/) — entrpi (independent) · Independent run · published 2 Jun 2026 · read 9 Oct 2026. Pass@1 for MiniMax M3 13.3%–16.8% across 2 settings, with cost per task, tokens. (2 readings)
35. [Artificial Analysis (Benchmarking Grok 4.7)](https://artificialanalysis.ai/articles/benchmarking-grok-4-7) — Artificial Analysis · Independent run · published 21 Sept 2026 · read 9 Oct 2026. Pass@1 for Grok 4.7 73% [xhigh]; Grok 4.6 65% [high]. (2 readings)
36. [Mercor APEX (mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/)](https://www.mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/) — Mercor · Independent run · published undated · read 9 Oct 2026. Pass@1 for Claude Opus 5.5 72.3% [max]; GPT-6.1 Sol 72.3% [max]; GPT-6 Astra 72% [max]; DeepSeek V4.1 Flash 71.7% [max]; Claude Opus 5 70.2%–71.4% across 2 settings; Gemini 3.8 Flash 71.4% [high]; and 43 more models, with confidence intervals. (52 readings)
37. [Benchmark review: DeepSWE v1.1](https://epoch.ai/benchmarks/deepswe/review) — Epoch AI · Independent audit · published 7 Sept 2026 · read 9 Oct 2026. Rates DeepSWE v1.1 flawed: grading defects in at least 23 of 113 tasks, most from hidden tests colliding with tests the agent wrote.
38. [envcheck: auditing DeepSWE](https://usetokenless.com/blog/envcheck) — Tokenless · Independent audit · published 30 Sept 2026 · read 9 Oct 2026. Documents 70 ways a submission can rewrite test outcomes, because the test harness sits inside the repository the agent edits.
39. [datacurve-ai/deep-swe issue #103](https://github.com/datacurve-ai/deep-swe/issues/103) — GitHub (community discussion) · Discussion · published 30 Sept 2026 · read 9 Oct 2026. Notes the board has published no new model since 3 Sep and quotes Datacurve's CEO saying the team is building the next generation of DeepSWE.
40. [AI Gateway model list and prices](https://ai-gateway.vercel.sh/v1/models) — Vercel · Price list · published undated · read 9 Oct 2026. Per-token list prices used to estimate cost where a source published tokens but no cost: Laguna S 2.1 at $0.09 per million input tokens and $0.18 per million output tokens.
