Normal view

Google updates Android Bench with new LLMs, but Gemini still lags behind

8 July 2026 at 16:39

Code generation is emerging as one of the most popular applications for large language models (LLMs), but not all agents are equally good at all development tasks. Google created a benchmark earlier this year to evaluate how LLMs perform in Android app development, and Android Bench is getting a big update today. The leaderboard now includes a raft of new models, and Google has adopted a new framework that should be easier to use. Developers are invited to run their own tests and submit feedback that could shape the future of Android Bench.

While they are popular coding tools, LLMs don't get everything right. Separating the useful outputs from straight-up slop means choosing the right tool. Android Bench aims to demonstrate which AI agents do best on a suite of 100 Android development tasks. After launching Android Bench in March, Google has added metrics like cost and efficiency, as well as open-weight models.

To keep Android Bench relevant, Google is updating the test with eight new models, including all the latest heavy-hitters: Claude Fable 5, Claude Sonnet 5, Claude Opus 4.8, GLM 5.2, Kimi K2.7 Code, MiniMax M3, Qwen 3.7 Plus, and Qwen 3.7 Max.

Read full article

Comments

© Ryan Whitwam

Secret Claude tracker shocks users after Anthropic’s anti-surveillance stance

Anthropic quickly removed a tracker secretly monitoring Claude Code users in China after a security researcher exposed the hidden code and condemned the spyware-like tracking as a “serious breach of user trust.”

Last week, a web developer known as “Thereallo” was researching privacy issues in Claude Code and was shocked to find that the AI firm was using “prompt steganography” to hide code that tracks Chinese users “in plain sight.” This code wasn’t malicious, but it was sending information to Anthropic that most users wouldn’t detect, relying on shorthand markers to quietly flag users’ timezone, proxy, and potential connection to Chinese AI labs that Anthropic has accused of distillation attacks.

On X, Anthropic engineer Thariq Shihipar confirmed that the tracker was added to Claude Code as an “experiment” in March. According to Shihipar, the code “was meant to prevent account abuse from unauthorized resellers and protect against distillation.” Regarding the former, The Washington Post found unauthorized retailers have sold access to free models for $1 a month, and pro subscriptions that can cost $100 monthly sell for "as little as $12."

Read full article

Comments

© SOPA Images / Contributor | LightRocket

Trump gets OpenAI to offer US 5% stake, far lower than Sanders’ target

OpenAI CEO Sam Altman is reportedly in active talks with the Trump administration about the US potentially acquiring a 5 percent stake in the leading AI firm.

Insider sources told the Financial Times that these talks are in “early stages,” but Altman “has argued that giving the public a financial stake in the company is the best way to share the upside of AI.”

Donald Trump favors the idea, and his administration has reportedly been talking to several AI firms about the possibility. According to FT’s sources, other companies approached to share similar stakes include Google and Meta.

Read full article

Comments

© NurPhoto / Contributor | NurPhoto

Musk’s X poses “serious risk to Americans’ privacy,” advocates warn FTC

Ahead of a July 2 deadline to submit public comments, advocates are warning the Federal Trade Commission that it must keep close watch over Elon Musk’s X and firmly reject a recent bid to end the agency’s ongoing audits of the platform’s data handling.

Last month, the FTC posted a notice explaining that X had argued that an FTC order was no longer necessary due to changes Musk had made to the platform.

The initial order came as a penalty after the FTC found that a coding error had caused then-Twitter to improperly share users’ contact information for ad targeting that had initially been submitted for two-factor authentication. Under the order, X is subjected to costly independent audits, and the FTC has authority to demand documents to ensure compliance with data privacy laws without taking additional legal action.

Read full article

Comments

© CHARLY TRIBALLEAU / Contributor | AFP

After spooking Trump into safety testing, Anthropic AI models get global release

The US has lifted export curbs on Anthropic’s newest Claude models, Fable 5 and Mythos 5, about three weeks after the Trump administration flagged the models as national security risks.

As of today, Anthropic confirmed in a blog post, Fable 5 will be available globally, and US organizations have had access restored to Mythos 5 since June 26. Anthropic said it is now working with the government to expand Mythos access to a “broader set of domestic and international partners in the Glasswing program.” That program allows cybersecurity researchers at trusted companies to access Mythos for defensive purposes.

In a letter to Anthropic viewed by Reuters and The New York Times, Commerce Secretary Howard Lutnick said Anthropic would “no longer need a license for exports or in-country transfers of its Claude Mythos and Claude Fable AI models.” The letter acknowledged that Anthropic had “taken steps in close coordination with the US government to address the risks” posed by the models.

Read full article

Comments

© NurPhoto / Contributor | NurPhoto

❌