Google is changing how it tests AI models for Android coding following new rankings led by Fable 5.
Google is changing how it tests AI models for Android coding following new rankings led by Fable 5.
Google released the “Android Bench” earlier in the year as a glanceable ranking system for the best AI models. Those rankings are all based around the model’s ability to code for Android – a general ranking system it is not, but incredibly useful for developers.
Google says there are a couple of changes coming to how it ranks those AI models for Android development. The core benchmarking system is being changed to the standardized Harbor framework. As it stood, models like GPT-5.5, Claude Opus 4.7, and Gemini 3.1 Pro Preview were ranked based on a mini-swe-agent v1 benchmark tool developed for general use.
Switching tracks and opening the system up to the Harbor framework allows Android developers to use the same tools to analyze AI models for individual use cases. In fact, Google says it’s opening up the Android Bench to users willing to submit Android development tasks. Those will be used to evaluate how models handle each scenario. Developers are also being invited to share their benchmark evaluations.
Android Bench adds Claude Fable 5From the beginning, we’ve valued an open and transparent approach, which is why we made our original methodology and test harness publicly available on
GitHub. You’ve asked for a way to provide feedback on our dataset, so now we’re taking collaboration a step further by giving you, the Android developer
community, a chance to shape Android Bench.
To top off the new approach, Google has refreshed the Android Bench list using the new framework for AI models. Each model between the closed-weight and open-weight variants has been reevaluated on the new testing bench.
Unsurprisingly, Claude Fable 5 sits at the top of the list with a rather comfortable lead. Google gave it a score of 84.5, a healthy 4 points ahead of GPT-5.5, ranked at 80.2. Claude Sonnet 5 is nearly 10 points lower than Fable 5. Anthropic’s power claims seem to be reasonable, and it’s worth noting that Fable 5 still has heavy restrictions put in place.
Most AI models that had a spot in the previous Android Bench rankings have received a new score, since the underlying rules have essentially changed. Here are the new rankings as Android Bench has them listed:
| Model | Score | Avg Latency | Avg Cost |
|---|---|---|---|
| Claude Fable 5 | 84.5 | 8.0 | $133.2 |
| GPT 5.5 | 80.2 | 15.7 | $138.3 |
| Claude Sonnet 5 | 76.2 | 12.3 | $99.9 |
| GPT 5.4 | 74.1 | 8.4 | $83.4 |
| Gemini 3.1 Pro Preview | 73.7 | 10.6 | $87.4 |
| Claude Opus 4.8 | 72.4 | 6.7 | $88.0 |
| GLM 5.2 | 72.2 | 38.9 | $117.0 |
| Gemini 3.5 Flash | 71.1 | 28.3 | $165.6 |
| Kimi K2.7 Code | 70.4 | 31.8 | $48.1 |
FTC: We use income earning auto affiliate links. More.
| # | Наименование новости | Тональность | Информативность | Дата публикации |
|---|---|---|---|---|
| 1 | Gemini 3.5 Flash lands on Google’s Android coding rankings, but it’s 3x the cost for slower performance | 0 | 13.22 | 12-06-2026 |
| 2 | Google изменила критерии отбора лучших ИИ для создания приложений под Android | 0 | 5 | 10-07-2026 |
| 3 | Claude Opus 5 launches with similar performance as Fable 5 for ‘half the price’ | 0 | 18.12 | 24-07-2026 |
| 4 | You can thank AI for Google Keeps’s new large FAB on Android | 0 | 3 | 04-06-2026 |
| 5 | Google will let websites opt-out of AI Mode & Overviews in Search | 0 | 5 | 03-06-2026 |
| 6 | Google Finance now available as dedicated Android app | 0 | 5 | 25-06-2026 |
| 7 | Google Is Cooking: Gemini 3.5 Pro Mogs Claude Fable 5 in Arena Leak | 0 | 0 | 02-07-2026 |
| 8 | Android developer verification on track for September, ‘Verifier’ service will soon auto-install | 0 | 5 | 18-06-2026 |
| 9 | Google makes AI Mode Pro visuals free ‘this summer,’ details Gemini tools for 2026 World Cup | 0 | 9.45 | 09-06-2026 |
| 10 | Fable 5 just set a new AI freelance work performance record - but it can't replace humans yet | 3 | 7 | 02-07-2026 |