Google has begun rolling out its long-awaited flagship artificial intelligence model, Gemini 4 Argon, but serious doubts are mounting within the tech giant. Employees note that despite impressive results in standard benchmarks, the new development demonstrates much more modest performance when solving applied tasks, especially in the field of computer programming.

Testing features and real-world performance

Google management presented the Gemini 4 Argon model to a narrow circle of cybersecurity partners and promised to expand access for paid tier subscribers. Official representatives emphasize that in a number of tests, the novelty outperformed competitors and demonstrated a high level of security. Nevertheless, sources within the company point to a significant gap between laboratory tests and real-world effectiveness, causing the model to struggle with frontend development and complex practical use cases.

Contradictory data

Attitudes towards Gemini 4's capabilities within Google are divided. Critics argue that the model lags behind competitor products like Anthropic and OpenAI in speed and code generation quality, while also criticizing it for its excessive size, which makes operation too expensive. At the same time, official management and some engineers insist that the model has undergone rigorous testing, refuting rumors of its inefficiency, and citing absolute trust in the DeepMind team.

Significance for the Google ecosystem and market risks

The success of Gemini 4 is critical for the company, as various modifications of the Gemini lineup are being integrated into all key products of the ecosystem, including Search, Maps, Gmail, and the Chrome browser. Experts suggest that Google may have gotten carried away with "benchmarking"—optimizing AI solely for tests at the expense of practical utility. Amid competition with AI agents from OpenAI and Anthropic, any delay or vulnerability in positioning new platforms could cost the company a significant market share.