Gemini 3.7 Flash targeted agent work per dollar
Gemini 3.7 Flash moved Google's Flash line up the agent-efficiency frontier. Artificial Analysis scored it four points above 3.6 Flash at 1.7 minutes per task and $0.40 under introductory pricing, while GitHub began rolling it into Copilot; however, Google's own benchmark gains remain vendor-reported and do not establish production reliability for all workloads.
Gemini 3.7 Flash is here. It’s stronger for coding, knowledge work, and web development. 🧵
Google has released Gemini 3.7 Flash, improving 4 points over Gemini 3.6 Flash and reaching the Intelligence vs. Time per Task Pareto frontier @GoogleDeepMind has released its third new Gemini Flash model in three months. Gemini 3.7 Flash (high) scores 56 on the Artificial…
Google has released Gemini 3.7 Flash, improving 4 points over Gemini 3.6 Flash and reaching the Intelligence vs. Time per Task Pareto frontier @GoogleDeepMind has released its third new Gemini Flash model in three months. Gemini 3.7 Flash (high) scores 56 on the Artificial Analysis Intelligence Index, just behind GPT-5.6 Terra (max, 57) and Muse Spark 1.2 (xhigh, 57) We benchmarked Gemini 3.7 Flash across all three reasoning levels (high, medium, low) ahead of release. With high reasoning, Gemini 3.7 Flash is a 4 point improvement over 3.6 Flash, while achieving an average Time per Task of 1.7, 40% faster than GPT-5.6 Terra (max). This places Gemini 3.7 Flash on the Intelligence vs. Time per Task Pareto frontier, reinforcing Google’s focus on speed across the Gemini Flash family of models Key benchmarking results across Gemini 3.7 Flash’s three reasoning levels: ➤ 4 point Intelligence Index improvement: Gemini 3.7 Flash (high) scores 56 on the Artificial Analysis Intelligence Index, up 4 points from Gemini 3.6 Flash. The improvement is driven primarily by gains on agentic evaluations, including Tau3 Banking (+3 points), Terminal-Bench v2.1 (+8 points), and GDPval-AA v2 (+103 Elo). With medium reasoning, Gemini 3.7 Flash scores 53, matching DeepSeek V4 Pro 0813 (max, 53) and GLM-5.2 (max, 53). With low reasoning, it scores 51, just behind DeepSeek V4 Flash 0731 (max, 52) ➤ Pareto frontier on Intelligence vs. Time per Task: Gemini 3.7 Flash produces ~340 output tokens per second, nearly 3x the output speed of GPT-5.6 Terra and GLM-5.2. With high reasoning, this translates to an average Time per Task of 1.7 minutes, placing Gemini 3.7 Flash on the Intelligence vs. Time per Task Pareto frontier ➤ 30% lower Cost per Task than Gemini 3.6 Flash: Gemini 3.7 Flash retains Gemini 3.6 Flash’s standard pricing of $1.50/$7.50 per 1M input/output tokens, however Google is offering discounted pricing through the end of the year at $0.75/$3.75 per 1M tokens. At this discounted price, Gemini 3.7 Flash (high) costs $0.40 per Intelligence Index task, 30% less than Gemini 3.6 Flash and matching Muse Spark 1.2 (xhigh, $0.40). With medium reasoning, Cost per Task falls to $0.26, placing the model on the Intelligence vs. Cost per Task Pareto frontier ➤ Leading performance on AutomationBench-AA and AA-AnalystAgent: On AA-AnalystAgent, our recently released benchmark measuring models’ ability to answer complex questions about spreadsheets and documents, Gemini 3.7 Flash (high) achieves the highest pass^5 score at 60%, ahead of Claude Opus 5 (max, 54%) and Fable 5 (49%). Gemini 3.7 Flash also leads AutomationBench-AA, our benchmark of agentic capabilities in simulated SaaS environments, with a score of 62.7%, ahead of Kimi K3 (max, 53%) and GPT-5.6 Sol (max, 51.2%) Key model details: ➤ Context Window: 1M tokens, unchanged from Gemini 3.6 Flash ➤ Multimodality: Text, image, video, and speech input, with text output ➤ Pricing: $1.50/$7.50 per 1M input/output tokens at standard pricing. Google is offering discounted pricing of $0.75/$3.75 per 1M tokens through the end of the year. Cached input tokens retain the same 90% discount
🆕 @GoogleAI's Gemini 3.7 Flash is now generally available and rolling out in GitHub Copilot. Early testing shows: ➡️ improvements in web and app development and agentic coding workflows ➡️ improvements in code quality, final-output presentation, codebase research, and…
🆕 @GoogleAI's Gemini 3.7 Flash is now generally available and rolling out in GitHub Copilot. Early testing shows: ➡️ improvements in web and app development and agentic coding workflows ➡️ improvements in code quality, final-output presentation, codebase research, and verification during complex coding tasks Try it out in the GitHub Copilot app, CLI, and @code. https://t.co/g8iY3a3RRC

The gains come in spurts as spending shifts to new chip generations. Performance per $ was nearly flat in 2023 as most spending stayed on the H100 throughout the year. It rose 44% in 2024, as the B200 began shipping, then jumped 80% in 2025 as the majority of spend shifted to Blackwell.
With each new generation, chipmakers sell far more powerful chips, often at higher prices. In 2025 dollars, the GB300 costs about 5.5 times the P100's 2016 launch price, yet it delivers roughly 200 times the performance, making it about 37 times more cost-effective.
The performance per $ spread across chips is large. Google’s TPU v6e delivers 6x the performance per dollar of an H100, partly because Google pays the lower margins of the suppliers that design and manufacture its TPUs, rather than the larger price premium Nvidia charges for its GPUs.