Google has released Gemini 3.6 Flash, replacing the short-lived 3.5 Flash model that debuted at I/O in May. The update offers improved coding performance, scoring 49 percent on the DeepSWE test compared to 3.5 Flash's 37 percent, and shows a slight gain in computer use capability, rising to 83 percent from 78.4 percent on the OSWorld benchmark. A key focus is efficiency—Gemini 3.6 Flash uses about 17 percent fewer tokens, reducing API costs to $1.50 per million input tokens and $7.50 per million output tokens, down from $9 for output. It now supports computer use as a standard feature in the Gemini API, aimed at improving accuracy and reducing steps in agentic workflows.

Alongside 3.6 Flash, Google introduced Gemini 3.5 Flash Lite and Gemini 3.5 Flash Cyber, the latter being the company's first AI model tailored for cybersecurity tasks. Flash Lite claims a speed of 350 tokens per second, positioning it as Google's most efficient modern AI, with pricing set at $0.30 per million input tokens and $2.50 per million output tokens—slightly higher than the previous 3.1 Flash Lite. Despite these releases, the anticipated Gemini 3.5 Pro, expected in June, has not launched. Google has not provided a release date. The company stated that changes in 3.6 Flash were made in response to user feedback, particularly around code generation shortcomings in the earlier model.

💡 NaijaBuzz Take

Google replaced a model it just launched weeks ago, suggesting the initial version did not meet internal or user expectations. The focus on token efficiency over raw performance indicates a shift toward cost management rather than capability gains. This may affect developers relying on consistent model stability for production systems.

Editorial note: AI-assisted opinion, not established fact. Full disclaimer →