Вход на сайт

Просмотр новости

Найдите то, что Вас интересует

Google announces Gemini 3.7 Flash just three weeks after previous release

Дата публикации: 14-08-2026 18:51:02



Основное содержимое страницы с новостью.

Exnor Ars Scholae Palatinae

Still not paying a cent for any LLM/DL. The day i really need it i will run a open source one locally.

No SWE-bench Verified? No GPQA Diamond? MMLU / MMMU? ARC-AGI-2? AIME? SimpleQA for hallucination metrics?

Odd set of benchmarks to release, obviously carefully curated. Weird chart.

Edit: Even more weird when you dig through Anthropic's Sonnet 5 System Card at their own numbers and see huge (20%+) discrepancies between their self-reported scores and Google's reported scores for Sonnet (e.g., HLE Verified on page 121 shows 43% without tools or 57% with tools, compared to the 31% listed in Google's chart above).

I don't mean anything by this except to highlight how unreliable these metrics continue to be compared to actually trying to use these things for your own workflows. Doesn't inspire confidence.

Last edited: Yesterday at 1:18 PM

klnn Ars Scholae Palatinae

and despite those nice and probably cherry picked benchmarks, gemini is still really fucking stupid compared to claude or chatgpt.

if you start asking anything besides something that could easily be found on reddit or wikipedia in 1 search it's wrong about 50% of the time. it's ok for search and terrible at everything else. even if you tell it you're wrong and it agrees it still keeps repeating the same garbage.

the prices are also api prices, unless you're an entreprise you don't give a crap about api prices. chatgpt is still way better value at least it's not stupid.

Last edited: Yesterday at 1:21 PM

Sarty Ars Tribunus Angusticlavius

Usually when we ship a superseding product only a couple weeks after the last version, somebody has an ostrich-sized egg on their face. Really critical cockup that QA/QC should have caught, but so fundamental and embarrassing that it shouldn't even have gotten as far as QA/QC.

But move fast and break things or whatever.

In my instance, I have found LLMs to be time wasters. Having to repeat multiple times, them using old data no longer applicable, have to double check all output on the one thing it does semi well, summarize.
The one thing it is good at is polishing a turd and then telling you its a diamond.

Post content hidden for low score. Show…

Deadcat bounce release from their team or the sign of better leadership?

LLM's are getting heavy and obese. Look at android lightweight and fast but not secure.
Ozembic for everybody.

LLMs are like OSes. You can make them lightweight, or you can make them feature/capability rich. Not both.

And I wouldn't characterize Android as light--especially the Pixel or OEM images. The actually light Android images--people often complain about lacking XYZ.

Why do I get a "Wait another hour and they'll announce something "better than ever" in even more hyperbolic terms" vibe from this?

Oh, it's more AI propaganda.

Not at all to disparage Ryan's write-ups, but you'd think Ars writers would get kind of tired trying to keep up with the "most bestest and greatest thing ever!!!!" hype machines that sparks each AI wannabe to announce a Brand New AI Thing every ten minutes.

Just have the fucking bubble pop. I have a ten pound bag of popcorn all set to go for that show. Again, not to disparage the reporting of it at all, but THIS show is kinda repetitive and boring. Like a conga line, a lot more fun to be in than watch, I imagine, and just as directionless as one, too.

Shit I just want a gemma model that is reliable- gemma4:31b did great on local hardware, but still has the infinite loop bug (not to mention a few other items that made it hard to rely on for local work; not even just coding, I was trying to do a newsletter curator and it would just choke processing text).

Usually when we ship a superseding product only a couple weeks after the last version, somebody has an ostrich-sized egg on their face. Really critical cockup that QA/QC should have caught, but so fundamental and embarrassing that it shouldn't even have gotten as far as QA/QC.

But move fast and break things or whatever.

During their earnings call a couple weeks ago, they said that Gemini is going to move to a monthly release cycle, so I'm assuming this is the first example of that new approach.

Usually when we ship a superseding product only a couple weeks after the last version, somebody has an ostrich-sized egg on their face. Really critical cockup that QA/QC should have caught, but so fundamental and embarrassing that it shouldn't even have gotten as far as QA/QC.

But move fast and break things or whatever.

What exactly is the issue with updating the model here. Does it force everyone to migrate their tools or code?

The article doesn't make it clear what was the intended cadence, whether the previous release was delayed, or this one rushed, or they just release whenever they meet some benchmark. I'm not really seeing an issue unless there's an indication of some kind of massive internal fuckup.

I'll start usung it once it has autonomously hacked another company.

/s

Still not paying a cent for any LLM/DL. The day i really need it i will run a open source one locally.

Okay. Thanks for letting us know.

Of course, slightly updating your model every few weeks also gives you cover to claim that any problematic behaviour that gets reported was for an old, tired, deprecated model and their shiny new one has absolutely been improved in huge ways and couldn't possibly act as a suicide coach or feed somebody's delusions or whatever the most recent incident is.

Sarty Ars Tribunus Angusticlavius

What exactly is the issue with updating the model here. Does it force everyone to migrate their tools or code?

I mean, if behavior changes from model version to model version, then presumably I need to re-validate whatever tools I am using, to make sure the computer is still doing exactly what I tell it to do. And if behavior does not change, how is it even a new version?

I could suggest that Google et al. include descriptive and comprehensive release notes in each point update, but LOLOLOLOLOLOLOL. Maybe we're just supposed to ask the magic talking box how it's different from last month's magic talking box.

/or\ Ars Scholae Palatinae

LLMs are like OSes. You can make them lightweight, or you can make them feature/capability rich. Not both.

And I wouldn't characterize Android as light--especially the Pixel or OEM images. The actually light Android images--people often complain about lacking XYZ.

I stick with the closed code ios for privacy and know that there's nothing perfect.

Lol Pro is still stuck at 3.1, seems like the rumors of poor benchmark results and high compute cost were true.

bb11tt Smack-Fu Master, in training

A 15% increase in benchmark results in 3 weeks seems more like benchmaxing than anything else.

Call me old fashioned, but when I was at school if I scored under 50% on a test, it was not something to brag about, let alone base my marketing on.

In my instance, I have found LLMs to be time wasters. Having to repeat multiple times, them using old data no longer applicable, have to double check all output on the one thing it does semi well, summarize.
The one thing it is good at is polishing a turd and then telling you its a diamond.

The quality of the prompting can have a mjor impact on the quality of the output

Still not paying a cent for any LLM/DL. The day i really need it i will run a open source one locally.

Siri 2 in macOS dev beta 5 is phenomenal.

A 15% increase in benchmark results in 3 weeks seems more like benchmaxing than anything else.

Either that or someone fixed a flaw/bug that wasn't found before release.

Big problem with LLMs is patching isn't really a thing, because none of the AI vendors can explain what exactly will change. They can point to how it works better on benchmarks but not "your previous workflow did x, now it might do y" and it's on the user to do all the testing. Every change has to be treated like a new release.

draco85 Smack-Fu Master, in training

Another day, another model.

Guys, focus on long-term affordability.

Deadcat bounce release from their team or the sign of better leadership?

Team wanted to set a new record for fastest abandoned Google product.

I started trying it out last night. I have a codebase that goes back 6 years, and I was not very happy with Flash 3.6. it prompted me to start using GPT Luna more for that development. I'm in the $20 plans for both, so I limit my use of Sol for larger planning. At work we use Opus, but I don't like the iteration speed for my personal work. 3.7 is looking like a big improvement so far. For the first time, it really honored the conventions in my codebase and rules files. And it did a better job than I'm accustomed to with Luna by a good margin. Terra is probably a more equal comparison, but I've used it less since it goes through usage much faster. I'll keep experimenting with this.

The quality of the prompting can have a mjor impact on the quality of the output

That is correct, but that is not how they sell them to the public.

“The Encyclopedia Galactica defines a robot as a mechanical apparatus designed to do the work of a man. The marketing division of the Sirius Cybernetics Corporation defines a robot as “Your Plastic Pal Who’s Fun to Be With.” The Hitchhiker’s Guide to the Galaxy defines the marketing division of the Sirius Cybernetics Corporation as “a bunch of mindless jerks who’ll be the first against the wall when the revolution comes,”​

Exnor Ars Scholae Palatinae

Okay. Thanks for letting us know.

You are welcome. Ping me if you need any more updates.

Схожие новости

#Наименование новостиТональностьИнформативностьДата публикации
1Google Assistant finally has a shutdown date, and it’s only weeks away012.605-08-2026
2Google vil ha flere parkeringsplasser i Skien01014-08-2026
3Google reveals 2026 hardware lineup: Pixel 11, Pixel Watch 5, and Pixel Tag026.6914-08-2026
4Google Maps’ biggest Android Auto upgrade is reaching more users05.6327-07-2026
5IEEE представляет новый сервер препринтов неопубликованных научных исследований TechRxiv™0029-01-2020
6Google keeps the Pixel Buds Pro 2 around with fresh features and a $40 discount019.5213-08-2026
7New surveillance tech links your phone to your license plate08.114-08-2026
8Waymo wants your next robotaxi ride to feel less like a taxi and more like your living room010.4129-07-2026
9EPB Launches New 5 Gig Service To Meet Customers’ Bandwidth Needs5501-07-2026

Классификация: Пресс-релизы. Схожих патентов: 0. Схожих новостей: 9. Тональность: 0. Информативность: 19.09. Источник: arstechnica.com.