AI

Google Rolls Out Gemini 3.8 Flash With Deeper Reasoning and Higher Token Consumption

The latest model matches prior per-token pricing but spends more compute on complex agentic workflows.

  • Google has officially launched Gemini 3.8 Flash, pitching the model as a more persistent reasoning engine designed to handle multi-step tasks and iterative tool calling.
  • Evaluation firm Artificial Analysis noted that while Gemini 3.8 Flash ranks as the most cost-effective model measured at its intelligence level, real-world task costs run approximately 40 percent h…
  • Google highlighted major gains across software engineering and autonomous task execution.
Google Rolls Out Gemini 3.8 Flash With Deeper Reasoning and Higher Token ConsumptionThe Scale Report

Google has officially launched Gemini 3.8 Flash, pitching the model as a more persistent reasoning engine designed to handle multi-step tasks and iterative tool calling. Arriving only weeks after Gemini 3.7 Flash, the new release retains identical base rates of $0.75 per million input tokens and $3.75 per million output tokens. However, Google cautioned that total operational expenses could climb because the model consumes additional tokens when running higher-effort reasoning workloads.

Evaluation firm Artificial Analysis noted that while Gemini 3.8 Flash ranks as the most cost-effective model measured at its intelligence level, real-world task costs run approximately 40 percent higher than 3.7 Flash. That jump is driven by a 30 percent increase in output volume alongside more interactive cycles during autonomous evaluations. Developers looking to preserve tighter budgets can still access Gemini 3.7 Flash, according to reporting by The Verge.

Google highlighted major gains across software engineering and autonomous task execution. The model led benchmarks on the DeepSWE v1.1 evaluation against Anthropic rivals like Fable 5, while also topping domain-specific tests including the Vals Finance Agent V2 and Harvey's Legal Agent benchmarks. John Ennis, chief executive of Aigora.ai, praised the model for delivering coding performance on par with Anthropic's Opus 5 at faster speeds and lower overhead.

Alongside the consumer rollout on Google AI Pro and Ultra subscriptions, Google introduced Gemini 3.8 Flash Cyber. The specialized variant is restricted to governments and roughly 650 vetted organizations, including CrowdStrike and the Center for Internet Security, under Google's Fairwind Program. The system integrates with Google's CodeMender agent to identify and patch security vulnerabilities across public infrastructure, supported by safety guardrails covering cyber offense and chemical, biological, radiological, and nuclear materials.

The swift cadence between Flash releases underscores how top AI developers are shifting from purely expanding model parameters to maximizing inference-time compute. By encouraging models to deliberate and chain internal tools before answering, Google delivers frontier-grade capabilities in lightweight packages. Yet this dynamic also shifts financial planning for enterprise developers, who must balance higher accuracy against the variable costs of verbose reasoning loops.

Reporting based on coverage from AI | The Verge.

The daily brief

The biggest stories in AI, venture, sports business and culture - once a day.

One short email from The Scale Report. No spam, unsubscribe any time.

Read next