Opus is still best in class for this, but it's worth noting how well Gemini 3.7 does vs a more comparable LLM price wise, which is Grok 4.6: https://html.non.io/neonRamenGrok4.6 . I thought Gemini would blow Grok out of the water (it generally has in the past), but Grok has really caught up.
Other thoughts: I really think Google has fallen behind here. Even as a high speed offering (this build took ~7min, which is pretty good!), it wont be able to claim dominance for long with cerebras announcing the Sol preview today: https://www.cerebras.ai/blog/accelerating-gpt-5-6-sol-ultraf... .
It's not a bad model by any means, but I just don't know what situation I'd reach for 3.7 Flash first for. Google really needs a differentiator, especially given how hard it is to get an API key from them. They can't be high friction and non-pareto.
Can you help me understand how it is hard to get an API key from Google? You just head on over to http://aistudio.google.com/api-keys and create a key... not any different from platform.openai.com?
Disclaimer: I work in Google so it might be that this link is not publicly well known
At a high level though, as a rule of thumb Google assumes that they're serving companies at Google scale first, and at a human scale second. For other companies it's the opposite. Generally what that means is the first experience you get with a Google product will route you through 8 different dashboards to set up ACLs before you've hired your 2nd employee.
Similar experience here, for what it's worth, though I didn't get as far as you. I basically just stopped and didn't bother - it was easier to go through OpenRouter than spend more energy on it.
Also:
> Google assumes that they're serving companies at Google scale first
So much this. I'm currently grandfathered in until the end of the year on Google's Search API, but the $35,000 they want to continue usage of my < 1000 personal searches per month, not going to happen. It has honestly been easier to use Anthropic to help me build my own search index & crawling infrastructure than deal with Google.
I think that's actually a very interesting insight that would be helpful for PMs on GCloud to take note of. As a single founder, setting up Google Cloud, it's like they start out by assuming you're bigco, forcing (I assume most) of their users into a arduous process of removing components they don't need.
Google AI Studio is one of Google's solutions to this problem, but in typical Google fashion, it's bolted-on without any clear connection in the ecosystem. If you're also using GCloud, it's hard to remember it's even there.
OpenAI's platform, by contrast, is streamlined, easy to use. With Google, I feel like I need to wade through the documentation first before even using the darn thing.
FWIW I definitely did not not need to do anything like that to generate a key via AI Studio. It was like three clicks to get the free tier key, later on enabling billing was a few more plus typing in credit card info.
yeah, "enabling billing" is a whole other ordeal. Like it all makes sense, it's not that hard, it's just a lot of extra clicks where the competition doesn't require that. You can blow off this feedback as the user whining, but it's real friction and the competition doesn't have that, so users are going to go elsewhere if they can.
You need to create a Google Cloud project to create an api key and when you try to create one you very often get error messages like:
“Failed to create project, The request is suspicious. Please try again” or “ You do not have permission to create a key in this project”. You can then navigate multiple screens in GCP to make it work but it’s a hassle compared to any other provider (OAI/Ant/OpenRouter or any of the Chinese labs).
I didn't have that issue back when I originally created my API keys a couple years ago. Just out of curiosity I switched to a different Google account that had never interacted with AI Studio and never used Google Cloud console.
It was literally two clicks, and didn't even leave the page: the dialog asked to create a project and type in a name, I did that, clicked submit and then it was selected as the default project. One more click and I had the free tier API key.
Not saying you didn't have that experience at the time, but personally I have had zero issues with AI Studio and consider it the most dead simple/fastest dev dashboard to get started compared to the others like OpenAI/Anthropic (thanks to Google's free tier that lets you skip billing setup annoyances just to play around with Gemini).
Despite what HN threads (that are also frequently confused and talking about GCP instead) portray as universal/widespread issues or the process being complex and time consuming somehow.
Last time I used the AI studio free tier it was limited to one or two requests, effectively useless. I think people are complaining its too hard to set up a paid API key (no reason you should make paying customers spend more than a few clicks and a minute of their time to pay you)
We have probably 10+ years old account with Google cloud etc. We recently had a production deployment, I went over to AI studio to get new keys and it kept failing saying "Failed to generate API key, The request is suspicious. Please try again" - It was through my standard browser, same geo-ip. And it just worked after 2 days.
Maybe things have changed but it was a big mess trying to getting an API key from Google as an individual a few years ago. Way too much conflicting documentation.
Eventually I gave up and run a few hundred million tokens (edit a few billion) through openrouter.ai using Gemini Flash 1.5 to Flash 2.5
Every since price increases on Flash 3.0 I've stopped using Gemini, too expensive for basic classification, sentiment detection, ocr etc.
As other posters said Google assumes you are some bigcorp trying to use their products. The Vertex versus AI studio confusion
did not help.
it worked for me okay when I needed it for myself in my personal account. But when I tried setup this for a company I spent almost a day solving lot of small puzzles in GCE like how to tell CEO that he have to connect billing account created for other purposes (and he not even remember at time that it exist) to new project and all other things that others talking about.
That is easy but I’ve also found myself in account setup dashboards that were obviously geared toward enterprise trying to set up access to tinker with something AI related. It might have been TTS but it’s been a little while and I can’t quite remember.
Oh it’s worse than that. There are places where you CAN set a limit. This seems great until you are at the center of a huge traffic spike because of good pr and so you try to change it to a larger number only to be told you need to wait 24 hours for the setting to change.
Biggest traffic day of the decade and our site was down because of google.
You prepay for tokens exactly like the OpenAI/Anthropic dev dashboards when using AI Studio, which the link above is pointing to not GCP, also there are project specific spend caps now.
I use LLMs rarely, and only for digging into subjects which I can't find enough information using search engines. I only tried Claude and Gemini, but Gemini both returns faster and higher quality information which I can use for more targeted digging myself.
Google being Google, their models tend to be better at finding, organizing and presenting information, from my experience.
Is it? I think waferscale might actually be cheaper per-token, it's just so many more tokens, and of course right now it's not a full buildout so the availability is limited as well. I'd imagine they'll be migrating to whichever inference method is least expensive, and I expect asics to be the ultimate answer.
Moving from either frontier intelligence or frontier latency to a single model that does both at the same time is potentially a game changer in certain industries. I can easily see e.g. hedge funds dropping tons of money on this, because it means they can now do the same thing as their competitors, but much faster. That's basically a license to print money.
There's a lot of trading that isn't proper "HFT", but where speed and latency still matter. Often you'll find this employed more as slippage reduction - i.e you're going to make the trade either way, but making it faster saves you a few bps.
I'm not sure what event-based traders are doing now, but back in the day NLP sentiment analysis was all the rage, so I'm assuming they've now incorporated LLMs too.
I am not sure this is the way to make AI more cost effective for such customers. If they are able to tweak any model for their use case it would be way more reliable and also way cheaper. In my opinion generic LLMs in the future will be just for attention economy or maybe government contracts. Everyone else will be running fine tuned free weight models or licenced closed source models (self hosted or managed).
I like using 3.5-flash-lite for doing cheap PDF and Image data extraction stuff. I don't think there is a better bang / buck model right now (3.1 is cheaper but a lot worse).
There is some irony being a developer and reading along the lines of: "oh look at the comparison between these models executing a task for a few cents on a job i'd be charging 1k minimum"
I don't think I'd say rarely. Companies rarely allocate the design resources to produce that, but the companies that do are typically much larger, so the actual number of individual developers that get rich mocks is probably closer to 40-50%.
I think both outputs are really good. I don't see a lot of differences. So what exactly should be looking at and notice that one model did worse or better than the other one.
EDIT: OKAY I see it's mostly the "image" generation, not so much the HTML... Noticeable in the food photos and the foodtruck/cart photo
How are you doing this with Opus. Clearly I’m missing something. I always turn to ChatGPT when I need images because Opus typically refuses. I’ve tried Claude Code and Claude online in the past. I’m pretty sure neither created images for me and I thought this was because Anthropic was focused on code.
I have a few custom ones (a post-trained flux 2 checkpoint for web design and a image->metalness map generator that the build step can call for more advanced lighting situations), but gpt-image-2 is better than my own for design, so it's weighted much more heavily in outputs my tool generates. I think gpt-image-2 currently generates 99%+ of the design outputs on diffui
I believe they are testing giving it an image, which you can do in Claude code by dragging/dropping into the terminal or copy/pasting, and asking it to build the html equivalent.
I'm curious how much the harness plays into this. I'm somewhat surprised by the gemini and grok results, they seem to have strongly deviated from the original images. I'm thinking maybe the harness has a big effect? It's possible to proxy in different models to claude code, if you're curious you might find it interesting to test!
I'm not sure what prompt you put in but did Gemini replace the all of the images in the original with its own? That would be really weird behavior unprompted.
The agent is told to generate assets as part of the buildout. It gets to decide what the prompt is for them / whether to do postprocessing like background removal / what type of asset to generate.
I think they're both testing very different things. The pelican test is testing if a LLM can come up with visuals on its own via writing bezier curves directly.
This is testing if it can match visuals that have already been established, and represent them with all the tools available to a web developer. The ramen example was chosen in particular because there are a lot of things that aren't easy to do with CSS, and require creative strategies: 45deg button cuts, angular repeating pattern elements, blending of raster art and svgs, microglyphs, low contrast subtle elements, etc.
Don't ask yourself whether it's a good design, as yourself whether it's a good test.
The "introductory pricing" for this 3.7 Flash model is really weird.
It's scheduled to double in price on December 31, 2026, but who would anticipate still using this model five months from now? Especially since 3.6 Flash came out just three weeks ago!
Then I ran it on high, medium and low thinking levels (oddly minimal is no longer an option, which WAS an option for 3.5 and 3.6) and got a pretty excellent pelican for the first two:
UPDATE: That was in Safari, but as pointed out in the replies here the pelicans do NOT render well in Firefox or Chrome! Best guess is that's because of this invalid filter in the SVG:
Filters are meant to contain additional elements, not be empty: https://drafts.csswg.org/filter-effects/#FilterElement - so maybe Chrome and Firefox remove the element that references the broken filter but Safari doesn't?
When a company gives away service a heavily subsidized service as a promo, the full cost of serving it (compute) can get classified as sales and marketing instead of just cost of revenue, which makes your gross margin look better!
At $work we still have some places using Opus 4.5. Even places using Qwen 2.5-VL, which is now 18 months old. It works, and upgrading is work (we'd have to validate the new model performs comparable in all the corner cases that currently work just fine)
Those kind of workloads would be hit by an end of introductory pricing. And it's exactly the kind of cases that are not very price sensitive. Where we are price sensitive we track new model releases closely, where we aren't other issues get priority as long as llm performance is good enough
For what it's worth, I use Brave and it didn't render properly in a normal tab, but rendered correctly on a private tab! Underlying engine is Chromium anyway, so make of that what you will.
Maybe it's a way to kick people off the old models. "If you still want to keep using the old model you can, but you'd better pay a premium."
I work for Google and their free, internal Gemini API isn't quite as graceful. They once turned down a model arbitrarily and it broke our tests. I had to scramble to fix it, then build warning systems for the turndown as well as a special validator to make sure any upgrades make the same determinations.
Yeah, so I'm going to confess. Ive built production AI systems that have very specific jobs, deployed them, and moved on. The cost is not noticeable, i had the system dialed in. Not worth the effort to reevaluate a newer model to see if its better in some way. Current model works, and I have other projects that are more important.
Ever since the insane discount with GPT-5.6 Luna, not much excites me anymore. I mean just look at the benchmarks, even though Gemini 3.7 Flash performs well on the DeepSWE 1.1, Luna (Max) still performs way better. I personally have stuck to Luna (Xhigh) because its been more than enough and does not bloat up the context window too fast with reasoning tokens.
GPT-5.6 Luna is an insanely powerful model for its price. It's been great for coding workflows where I guide the LLM's hand step by step. It's also insane to see my weekly limit drop by than 2% after an hour of coding ever since the discount.
However, I've noticed 2 drawbacks with Luna. Context rot is much more palpable than Terra and Sol. It tends to get confused and go into rabbit holes when it's context gets filled up. In addition, when instructions are vague, it performs poorly and tends to write way to more code than necessary, but that is to be expected of smaller models. In all, for clearly defined, bite-sized coding tasks, Luna's price-to-performance has been insane. It might have very well commanded the price tag of Sol if it came out just a year ago.
Yes it is cheap, but per task DeepSeek v4 Flash is a bit more expensive and lands between Terra and Gemini 3.6 Flash in quality. Closer to Gemini than Terra...
luna is the first model that has outdone gpt-5-mini on the pareto frontier for some of my high value, cost sensitive ai product workflows. it's both cheaper (by about 60% in real world use) and higher quality based on my test harnesses. I was really worried that costs would go up since there wasn't a replacement as of a few weeks ago and gpt-5-mini is scheduled to be sunset toward the end of the year. So long as they don't randomly sunset this model anytime soon, that worry has now subsided.
I am in the same boat as you. I am using Luna and DeepSeek Flash. Both super fast, super cheap, and I have not felt need for anything more capable in few weeks.
I'm curious how much people are manually curating context these days; I'm increasingly feeling for myself that it being auto-managed inside a front-end like claude code is not ideal, and I'd rather have more control over what exact files and pieces of discovery go into a particular prompt, and the ability to more easily "fork" a session and ask asides or make notes/todos in a way that doesn't disrupt or confuse a more focused task going on.
I don't think I want a gastown-style "just yolo everything" approach, in fact I really want more control over how decisions are made and with what info. Does this exist?
I am not sure this enlightens you with anything but I have a TODO.md file with three headlines. Todo, Doing and Done. The agent is aware of it and knows on which task we are on.
On complete, it moves the user story from Doing to Done. I also have a MEMORY.md file that the agent read and writes in the beginning of a new conversation and at the end of our conversation to update stale information. These files are referred to every time I start a new conversation.
Regarding forking, I know Codex has such button underneath each message that lets you fork the whole conversation. I usually do that when I want to sidetrack and discuss something.
I use no SKILLS or commands like /goal. I’ve come a long way with just prompts and markdown files. Its all a different way of encapsulating instructions anyways.
It’s a big difference but Luna is very usable. I’ve plugged it into the slot I used to have GLM 5.2 in; I think it’s just as good. And it is less costly. I have Sol do planning and design but do most task execution with Luna now.
what do you mean by reset at every turn? context stays until compaction. if you remove the reasoning tokens after every turn you will be constantly blowing cache which is far worse than filling up context.
Does it reset at every turn? From my experience in Codex for example, Luna (Max) fills the 256k token window relatively quick. The only thing lowering the context window again is the compaction.
They need to release benchmarks against Luna/Terra. Luna is much cheaper which feels like it undercuts the need for Flash.
I've always considered the Flash series of models to be for low-cost, high-volume, mostly text-based use cases (e.g. summarization, parsing, formatting), emphasis on low-cost.
flash-lite is more of their luna tier competitor but even still not quite there yet, but gemini's dominance on multimodal and image understanding i think really gets downplayed on this site when most people think the only think you can do with LLMs is write code
Ultimately it would track that in the real world, people will want to point cameras at things and get answers.
I pay for ChatGPT and Gemini, and while Sol is a total beast with anything text, it still poisoned my cucumber bed. Which I will be bitter about for at least a few years while the bed recovers. Gemini (even flash) is exceptionally talented at viewing photos and telling you what to do/what it is (and telling me I just misidentified the problem with my cucumbers and spraying off the "bugs" actually just spread the bacteria everywhere.)
That's been my association as well. I see Flash get brought up a lot in relation to things like OCR and PDF processing frequently, and a lot of other routine multimodal workloads.
Yes this is my impression as well. To be fair I didn't compare to Luna yet, but Gemini 3.5 Lite is a very good and cheap multi-modal data extraction model.
I wonder why comparison table doesn't include deepseek-v4-pro. In DeepSWE 1.1 DeepSeek is slightly below Gemini Flash 3.7, but it is much cheaper, and it is above Claude Sonnet 5 which was included.
Yes, was going to say I use it exclusively for video and audio. The ability to give it a YouTube link through the API and ask questions about it is awesome
For what it's worth the source of the data[0] does have 3.7 flash with all 3 reasoning levels. 3.5/3.6 are in fact just the single points though (high reasoning). The datapoint in the announcement screenshot is either med or high, but they're pretty much exactly the same so can't say for certain.
That graph has to be made because an Exec didn't like that graph went down to the right instead of up and to the right. How do you make a graph with 0 on the far right and counts up by going left of 0? What number line is that?
So it's better than 3.6 Flash, at half the price. I've been pretty excited about Gemini models recently, they just feel so fast after spending most of the day at work waiting for Opus 5.
Yes, 3.6 Flash is very fast. I used to get a fair amount of usage of the Gemini Flash models on the free tier. I signed up for their $4.99/month tier (includes 400GB of Google space which was also enticing) and it turns out I only get about 15 to 20 minutes of usage before I get a come-back-in-7-days message. Comically low usage limits on that plan.
It's funny that they don't mention this at all in the marketing or tech specs when it's obviously the biggest selling point by far. Without this it would be completely irrelevant.
Worth noting that OpenAI just announced that they got the full GPT 5.6 Sol model running on Cerebras at 750 tokens per second. No announcement of the pricing though...
> It's funny that they don't mention this at all in the marketing or tech specs when it's obviously the biggest selling point by far. Without this it would be completely irrelevant.
Good catch! You're right to point that out. My previous marketing copy missed that specific detail. Thank you for bringing it up!
Cerebras is crazy to watch on GPT OSS or Gemma, I feel like we need a new VibeOS demo but with Cerebras, the OS would literally build itself in a few seconds.
I use this in a customer facing application and Gemini’s speed makes the experience feel much better.
The application isn’t so complicated that you need opus level reasoning or code writing, we need “good enough” data retrieval and processing with natural language queries and the ability to answer follow up questions.
I've blown away by flash 3.6's speed while Opus chugs along for _hours_ on similar tasks. I've gotten into a opus designed -> gemini implemented -> opus reviewed dev cycle recently.
I am actively using Gemini flash to "translate" what Opus says into human language. I let opus do the design (with my assistance) and implementation, but then the report that Opus writes gets translated by Gemini so that I don't have to waste time to understand it.
You can also customize Gemini Flash. It's a niche thing benefitting few, but you can tune gemini-3.7-flash in Google Vertex (now named "Agent Platform"?)
Sol high is almost the same speed if you take into account drastically lower token use. Look at the artificial analysis speed vs token use. Gemini is 7x faster but 5x more tokens. And that's with Sol high being a substantially better model.
Edit: and Sol medium actually has the same AA intelligence score as Gemini 3.7, and has >7x fewer tokens, actually making it faster
is presumes you are doing longer difficult agentic tasks, if youre doing a simple problem in 1 or 2 shots, not really multi turn then theres no comparison.
Yeah we use it for auto-triage of incidents, attempts to auto-remediate, and escalation to human. But for actual development, it’s not a viable option for us.
Yup. I use it for a ton of mundane queries (stuff that I might have used Google search for in the past) and it's great. Nice and fast and correct more often than not, especially if you prompt it in a way that it invokes Google search (but filters out ads and SEO slop). It's even alright at programming tasks but if it stumbles then I'll escalate to Gemini Pro with extended thinking.
The multimodal abilities are great, but if you deal with text only, what is the benefit of using this over DS V4 Flash/Pro? 13-26x cheaper with comparable intelligence, and available across many different inference providers.
I fail to see the usecase where DS V4 Pro is not enough, but Flash 3.7 is - except multimodal.
Luna is similar, and also 8x cheaper. Source: artificialanalysis
The only benefit I can see is the speed, that looks to be outstanding, probably thanks to their TPUs.
for non-coding applications, i think speed is a real differentiator. Im building an app that uses LLMs for some functionality that the user would not have any reason to expect is using AI and therefore having then wait seconds or minutes is just not feasible. latency is a huge upside for me
Anything interacting with the real world seems like latency would be hugely important. Something more asynchronous friendly (like coding) is for obvious reasons over represented here
> 13-26x cheaper with comparable intelligence, and available across many different inference providers.
Well, compared to 2 months ago, it's no longer 100x more expensive for similar levels of quality...
If they continue monthly-ish releases by 3.9 - by Halloween - they should be close to the best in terms of what you get for what you pay for.
In 2 months, they've gone from basically the bottom of the pack to at least being somewhat usable and competitive.
OpenAI and Anthropic release in a month, and change things. OpenAI is claiming to be close to an Astra release - but that seems like a Fable type release - where they're just releasing a better more expensive model, not more cost effective models.
did you try to ingest 1M documents per hour with any provider except GCP with Flash? None work at scale. Deepseek, Luna, Mistral all fail. 1 in 3 requests is a fail. I stopped trying.
The only thing that works at scale is gemini flash.
I guess the question then becomes "are you sure you'll do text only?"
I could probably do text only for my workflow (feature development/debugging for web microservices) but sometimes it is easier to just toss a screenshot into the Claude prompt, so that gives it an edge.
If your workflow is 100%, certifiably never ever going to involve an image, then yeah, this isn't going to be huge.
The Gemini Flash models makes perfect sense to me coming from a company like Google. Google AI Mode for search is a product I really find useful. It makes sense that Google focused on smaller, faster yet smart enough models that wouldn't break your bank on inference. It plays well into their product ecosystem.
Google AI Mode consistently gets me consistently good results and good speeds. It really changes what "googling" is for me.
I similarly have a weird affinity for Gemini that I can't really articulate. I used Gemini's free chat and found it great for exploring technical topics (and random one-off general walking-around-questions) and appreciated its speed, tone and accuracy. I spent a month playing with Gemini CLI / Antigravity and found it also an effective coding agent, at least for my workflow (entirely in the loop development and review). I also was really surprised that I could just paste it images of a project I was working on and have it immediately understand what it was looking at -- which I've come to learn is considered a unique strong point for Gemini. I've been playing with GPT5.6 for about a month and it's definitely powerful but I honestly think I'll go back to Gemini. There's something kind of charming about working with an AI that not only is particularly good at web search and information gathering, but also one that doesn't feel like some superhuman overengineering freak when it comes to code.
I like Gemini (I'm just a dumb person without knowledge of 'benchmarks' or how x compares to y) while understanding that shoving it into Google auto summaries has been a bad idea and produces inaccurate results
> Coding and agentic tasks: Significantly higher quality on real-world software engineering and agentic benchmarks, improving issue resolution and reducing failed agent loops.
> Web development and stronger design parity: Generates higher-fidelity desktop and web application code directly from design mocks, with strong gains in design adherence and in auditing existing codebases against mocks to verify 1:1 design parity.
> Promotional pricing: Gemini 3.7 Flash will be available at an introductory price of $0.75/1M input tokens and $3.75/1M output tokens. We’re also applying this new rate to 3.6 Flash. Introductory pricing expires on December 31, 2026; after, $1.50/1M input tokens and $7.50/1M output tokens will apply.
Still no sign of 3.5 Pro. Will have to test it, low expectations given every other model from the Gemini 3 lineage, but one can hope. Just struggle to understand the promotional pricing being temporary for four months. Given this industry, I'd be hard pressed if 3.7 Flash was still in use by end of year, so why not make it the official pricing?
>Given this industry, I'd be hard pressed if 3.7 Flash was still in use by end of year, so why not make it the official pricing
It was probably to placate some kind of general internal pricing/revenue benchmark that doesn't account for new model releases. Politicians do shit like this incessantly and it reeks of bureaucracy.
I suspect it's a bit of a signal to investors etc.
"Hey, we are not in a race to the bottom. This is our usual pricing, but this now is a promotion because we know we're coming from behind and need to entice users."
They're drawing a line in the sand on monetisation and signalling that to everyone, while in reality offering it a deep discount (no idea if profitable or not) knowing that this model will probably be obsolete before then.
And I'm over here on SiliconFlow using Stepfun AI Step-3.5-Flash at $0.10/M input and $0.30/M output tokens (262K context window) for complex market analysis work in rust utilizing vectorized instruction sets. It provides me with amazing results.
I honestly wonder how long this calliope can keep playing before it crashes to the ground.
(I have no business relationship to anything mentioned here except as a regular retail customer who went bargain-hunting)
I don't get it, Google could heavily subsidy their Gemini models to make it more attractive, but they prefer to not do it. I don't know one soul who is using Gemini models to code.
Even OpenAI who doesn't have money or capacity is offering their Luna model at $1.2 per 1M/out.
Why would they? Unless they have lots of unused tpu real estate that they could host it on “for free” they would be bumping more profitable workloads off of machines to give away that capacity to people with zero long term loyalty. There is no business reason for google to subsidize these models.
OpenAI has too much money. They’re spending their money in stupid ways.
What you're not seeing are the subsidized Google Cloud startup credits, which includes Gemini. If you're in that program, you choose Gemini because it's essentially "free" and consistent.
It makes sense. All these models are money losing businesses.
As a business Google might want to focus on fundamental research 2-3 years from now and not compete on who acquires more money losing customers. Just stay little behind and invest money better.
On a related note, I see all these quantitative benchmarks and the models getting really good at them over time. One thing I've been wondering: if the GPT series of models performs so well quantitatively, why do I still kind of hate using them relative to Claude? There’s a missing “vibes” or “taste” benchmark I think.
This is genuinely a competitive model, considering it beats Claude Sonnet 5 on almost all benchmarks and is more than half its price. Seems like Google is back in the game, though not leading the frontier anymore.
Sonnet 5 is arguably the most cost ineffective model to ever be released, so that's not really impressive.
It can regularly cost more than Fable, take longer, and deliver far far lower quality.
I'm much more interested how this compares to Luna - which on price is terribly - but at least on quality the benchmarks make this look competitive / usable.
If Google continues monthly Flash releases like Sundar said they would, and they continue to have this much of an improvement in cost/quality - then in a few months this could reasonably be very competitive with the best of the best.
It is not there yet, but at least it's super fast, I guess.
Google is not currently in the lead for maximum model capability, but it is still very competitive (or even best) in the multidimensional capability, cost, and speed frontier.
Have you tried the new DeepSeek Pro v4, Qwen 3.8, Gemini 3.7 Flash, and Grok 4.6? Do they make sense for any use cases?
I'm currently using omp with Kimi K3 as the planner and DeepSeek v4 Flash 0731 as the implementer, or CC + Fable for planning and Opus 4.8 for implementation. For API(not coding), I just use DeepSeek v4 flash 0731 and MiMo.
I'm pretty happy where I am, but I'm wondering if these new models provide some new kind of advantage
It's basically a "if we really have to support this for a long time, we want to be compensated for that" pricing strategy. It's about long term maintenance cost being greater _because_ it will be irrelevant.
I was going to cancel my gemini membership today ..... still going ahead. In my experience, gemini 3.1 pro, 3.5, 3.6 flash constantly lie too much about completing their tasks whereas sol (even though equally dumb) never claims something has been done when it hasn't been.
i just wish google cloud ux was remotely as good as their models. they made some progress with their studio, but then in a true google fashion, product names keep changing (Gemini, Anti-gravity, Vertex, Google AI,...) as well as confusion and complexity for something as simple as registering agy cli with a Google cloud project.
today i wanted to link agy to a google cloud project, for that i had to enable 5 different APIs in google cloud UI, then create a subscription for Gemini Enterprise (whatever that is), then link it to a project, then assign it to a user. and after all that, agy couldn't find the subscription.
the best part: i couldn't cancel the subscription. so i just paid $35 for one month and left it.
Maybe this is just my experience, but have people had trouble with 3.6 Flash just... getting things it has seen in its context correct? I don't know if it's been insanely benchmaxxed or what, but it'll pull information from websites and immediately get it wrong the token after. Or for example (this is something that happened like yesterday) I asked it to compare the uses of A and B in a language I was learning, and the way I typed it was "Please compare how these two are compared differently: A VS B", and then... it proceeded to compare "VS" and "B". I'm not kidding.
Personally whenever I use Gemini I've just been using 3.1 Pro because I've had insane trouble with them getting things incorrect like this. Hopefully they'll fix it soon / they've fixed it with 3.7 Flash.
Artificial Analysis shows Grok 4.6 taking $1,068 to run their suite while Gemini 3.7 Flash takes $485. So it looks like Gemini 3.7 Flash is less than half the price in the real world.
Per-token cost isn't a great metric given that some use way more tokens than others.
> Hard to understand why anyone would choose 3.7 Flash under these conditions.. is Deepmind still a frontier lab?
At this point, I think they're mostly targeting Google One and Workspace subscribers, except doing worse compared to Microsoft because they don't have Microsoft's huge enterprise moat built from their DOS and Windows days.
Agree, with you but I'm still using 3.6 Flash because of tok/s/ latency/ uptime with high context. Tried Grok 4.6 and it was scoring lower on some internal benchmarks or slower.
It's on Google AI Studio, which I use for free when I'm not on computers I control.
It did fine on my usual benchmark about configuring old Sparc hardware, maybe output slightly faster than before. Even included something new to check in the firmware.
3.7 Flash gets 56 on AA up from 52 for 3.6 Flash. But it seems like this is at the cost of more output tokens per task: 3.6 Flash is 26k, 3.7 Flash is 37k. Due to 3.7 Flash's 2x slashed pricing it's still cheaper per task.
I want to like Gemini models but my problem thus far has been a lack of coding chops. They still make mistakes, importantly, without correcting them for things like hallucinated API calls or code that doesn't run but they never bothered building or running. I know a lot of this can be fixed with workflows but it still feels like a failing.
GPT-5.6 or Claude models haven't delivered to me non-running code in ages.
Whenever I have Gemini in the flow, it's fast, but mistake riddled. I have low confidence in the output.
I've had some success with Opus driving Gemini models. It's pointless for GPT family since Sol is cheap enough or can drive terra/luna for arguably better performance, same speed, and better outcome.
As for all of the talk in this thread about modalities. Every SOTA model takes screenshots and verifies work now. Grok-4.6 does this, Luna does it, etc. They can also all work _from_ a screen shot or mockup provided.
I don't think it's a major selling point when every model can do it well and reasonably fast.
That said, eagerly awaiting "pro" and improvements to antigravity.
It depends on if you are relying on all single shot tasks or are willing to iterate. Flash is quick and can make dumb mistakes, but it also can fix them quickly.
I've gotten good results with it, but it definitely is more hands on.
I don’t think the skill set required to write the Vaswani paper is the same as training and shipping frontier models like Gemini so I am not sure why people keep bringing this up.
Does Google believe people want fast models because they have some sort of evidence of that preference? Or are they no longer capable of delivering a Pro model?
I have read that "pro"/"opus"/etc models can actually be worse for everyday coding as they reason "too deeply" and turn over too many stones over-thinking the problem and potentially getting distracted.
This feels absurd to me (my gut is "I want the SMARTEST model I can get!!"), but often I find that my experience of using a flash/sonnet model for every-day workhorse coding they are better.
Its not the same thing, but when I think of that I am reminded of working with some engineers in the past who are incredibly smart and have PhDs (or to put it another way, over-qualified) and they were crap engineers because they'd just not be able to focus on the task and ONLY the task at hand and would get easily distracted by the "why" or "more interesting" things when I just asked them to fix a simple bug or whatever. Again, its not the same thing at all, but it certainly comes to mind when I think of this or experience a pro/opus model suggesting we make huge refactors when a tactical fix is all that is required etc.
Of course, the opus-sized models are great when it comes to huge comprehension/research/debugging efforts where the deeper reasoning is actually useful.
I find Fable completely unbeatable for anything code-related. It's the only frontier model that seems to come with sane defaults.
If it implements something simple like a file export, it just knows that the file should have a meaningful name. Vibecoded feature beats most software's lazy "untitled.png".
So, yes, I want the smartest model even for simple stuff. Maybe especially for simple stuff because the tokens burned will be trivial so the cost doesn't give lower models a comparative advantage.
True, if you have a codebase that works in practice but has dozens of loose ends and poorly defined edge cases than it can chase off into rabbit holes because "oh wait, what if x is undefined instead of null? How is y defined? This outdated package has long known severe security holes and should not be used anymore, do we actually need it?".
Probably both. Having a strong frontier model is necessary not just for the model itself but because it provides a halo effect for your entire line. So if Google could deliver a pro model they would. But I also think Google is targeting the wider market and not picking verticals like Anthropic does. A good enough model is good enough for most generalist tasks, and being fast and cheap is more important to less sophisticated users. Also can't forget Google is at every level of the AI vertical. They're not losing sleep because they're not competitive at the one level in which open weight models come out with the quickness. It reflects poorly on them, and from a marketing perspective its not good but in some ways its actually the least valuable place to be.
If you think about Google and their business/reach, fast and light models suite them the best.
Google probably crunches more tokens daily than the other labs combined, just because basically the entire global population uses Google (sans china) and Google has shoved Gemini into everything.
throughout history, Google has been obsessed with speed as a feature. that was a huge reason people used google search, and then chrome in the first place, and it think its really underestimated by people. Jeff Dean specifcally seems to think about this alot.
That would be concerning if true, since they seem to have made a heavy bet on multi-modal as the way forward.
I wonder if this counts as evidence against that hypothesis? That multi-modal is struggling to keep up with SotA and the best they can offer is competent and fast?
> * For 3.6 and 3.7 Flash, introductory price expires on December 31, 2026. Starting January 1, 2027, $1.50/1M input tokens and $7.50/1M output tokens will apply.
this is hilarious. it is not 2025 any more, by Jan 2027 there will be at least 3 newer generation of models (from other provider) released already. nobody would use flash 3.7 at that time.
sure we used to cling to gemini models in the past, demanding 2.5 models to continue to serve, but since google betrayed us with those price hike, people already spent their time making their production pipeline less dependent on google since then.
heck, even now I'm not sure I even care if they cut the pricing even lower. there are too many models with cheaper price and similar performance now.
They should call it 'face saving pricing after we realized just how terribly did we mis-price the flash 3.5'
> since google betrayed us with those price hike, people already spent their time making their production pipeline less dependent on google since then.
This is my first hand experience. I spent at least $3000 on gemini-3-flash-preview. And exactly $0 total on (3.5+3.6+3.7)
Maybe the business model is to break even on bleeding edge models while making money on the long tail of usage once systems are tuned for a specific model and running in production.
Think of it this way: you are at an enterprise business. You have a workflow implemented a year ago that is working just fine. Swapping out the model for a new one changes behavior in unpredictable ways. Eventually, you'll do it once cost is low enough, but it takes serious labor to validate this, so you'll wait a long enough time for Google to make money.
Prompts can certainly be tuned to a particular model, where updating the model actually results in worse performance. This is perhaps less true today than a year or two ago, but we have seen this on newer models as well. Typically, the less specific the instructions are, the less it's a problem. But sometimes you really need to get into specifics to get good results. Area of work is code porting and translation.
this is becoming less true with every generation of model.
a decent model with a decent harness will determine when the knowledge base is lacking and attempt to fill the holes; thus the good general models can be very easily brought up to speed on niche domains.
Some model are more aggressive by default, some are more verbose by default.
To get the result you want for your specific application, you run experiment with prompts and parameters.
it is partly true, but like I said it is not 2025 anymore. models now get released more often, and still have notable progress so they can safely replace the old models while being faster/cheaper. and thank to chinese models the pricing is pretty much stable and affordable now.
and now we have ai agents to automatic migrate the system with new models. in the past we would need to spend hours to design the prompts, then test the output, then write codes to babysitting it. nowadays any ai agent can do it effortlessly.
I think it's meant to make fun of the fact that Google raised prices on their models and people were upset, and this is Google's way of lowering back the price because by Jan 1st 2027, this model isn't going to be used since people will move on to the latest models.
Personally, I feel like Google blundered on their pricing because while I was using the free version of the Gemini harness, they took away most of the free limits and made people move over to their Anti-Gravity harness for no apparent reason. I was about to splurge for a Pro sub since I already used Google for extra storage but putting up limits like they did made me not want to trust they wouldn't do more price shenanigans. Now their models are behind and it seems like they're scrambling.
Nobody except corporations who built workflows on top of it and don't care about the price because the developer already moved on and nobody wants to touch it.
>3.7 Flash is available through the end of the year at an introductory price
1 of $0.75/1M input tokens and $3.75/1M output tokens. This price combined with the enhanced model performance enables developers and customers to scale production-ready agents cost effectively.
Introductory pricing until December 2026 implies no significant Gemini Flash developments until the next year.
I think it's just meant to make it more competitive, Gemini has kinda been behind in everything except maybe multimodal. It's only 3 weeks after Flash 3.6, so if they really wanted to, they could probably do a 3.8 Flash before then.
3.5 Pro was supposed to be around the corner two months ago. 4.0 Pro is some ways out as they recently stated they are seeing some promising early results from training. It didn't sound like a release is imminent.
Has anyone noticed that antigravity has been working really well for the last few weeks. Now with this model it should be working much better. Hope the Google AI Pro Subscription can be used to do some real agentic coding now.
So at this point new models seem to only care about one task, software development. This really was not the original pitch of ai and I do not see how it justifies the insane spend or valuations it has produced.
Think of the potential layoffs of highly paid employees!
But, I think it’s also based on what they are being used for, most LLM users are still mainly SWEs or similar as I understand and there’s a ton of data to train them for coding.
I agree its what they are being used for and their primary revenue source.
My point is mainly that was never the pitch that got ai the hype it did and imo doesn't justify the valuations even if we all lose our jobs to ai. Because it no longer seems like they even think its making other jobs go away.
I just use web chat as "harness"(lol) or interface and I have mostly switched to Gemini as the free limit basically never run out for me unlike ChatGPT and Claude.
So basically Google is 5 weeks behind with a Fast model that is as good (on the bench they picked it's mostly ahead btw) as the models that the two darlings of HN (OpenAI and Anthropic) released five weeks ago.
And yet the entire thread here is people bitching that it's neither 5.6-sol nor Opus or Fable 5.
BTW why are OpenAI and Anthropic even releasing models like terra/luna and Sonnet?
Why? Just why?
Is there a... market?
For you can't have it both ways: either Sonnet and terra/luna make zero sense for Anthropic and OpenAI or Google is a player.
Did the company fix the high friction between any service and their models' API?
I hope so. It seems mind boggling to me that an user needs to surf around different sections (plural) of google cloud console, then this Vertex and do a dozen clicks to issue a simple key.
You can use the Gemini API which is independent of the more complex Vertex AI API. Not sure whether you still have to visit the Google Cloud UI for some things (like billing) though.
Well, you do get 1 million tokens and the ability to reason over video natively and many of us are forced to pay for 20usd plan anyway due to google drive 5TB, not to mention notebooklm, so it’s not a nothing burguer, it’s just an almost nothing burguer
At the discounted rates, upgrading from 3 Flash to 3.7 Flash is finally reasonable.
In my evals 3.6 Flash (pre price change) was usually a bit more token efficient than 3 Flash, so I‘m expecting same or even lower cost-per-task on 3.7.
Offering a 'temporary introductory discount' until Dec 2026 on an LLM is hilarious. In this market, by Jan 2027 this model will be superseded by 5 different providers offering 10x the performance at half the post-discount price anyway.
I wasn't aware of this. Seems Google is lagging the big 3 (Anthropic, xAI, OpenAI) when it comes to frontier models for programming and hard problem solving.
I guess Google's betting on consumers being price-elastic (preferring to tradeoff intelligence for significant cost savings)
"Lagging" is putting it rather lightly. GDM is no longer a frontier lab.
FWIW neither is xAI, there is no "big 3". xAI has had momentary peaks (I think they are having one right now) but they have never been able to claim to consistently push the frontier in any particular direction. You can also infer they aren't a frontier lab from the fact that they sell their compute.
After being stuck with using GPT-5.6 models for the past few weeks, I have renewed faith in Google and everyone but OpenAI. The GPT-5.6 models are quite obviously benchmarkmaxxed to make they seem like they are intelligent but they are quite dumb outside anything that not a benchmarked task.
I also think Google is still the best at fitting the most overall intelligences into their models, but for some reason it seems like the model architecture is just bad.
Like, I understand everything, but by this time I don't give anything about any of those announcements.
Theoretically there is some difference between Fable and Opus or Grok and GPT, but at the end of the day I'd look at the bottom left of my screen and to my amusement find out that for the past 3-4 hours I've been using model ______.
If the results are semi-decent, I'd keep it on, if not - I'd randomly switch the model and try again.
Actual thing that would affect my selection would be a number of unused tokens I have left for a model ____ for this week.
Maybe it's cause I'm using those for programming and log parsing and all of them are decent enough, but other than that - there are no leaps I see.
IMO, they should drop their previous model (3.6 Flash) from the benchmark charts. I don't care how better this is compared with their previous model. What matters (to me) is:
1. How the new model performs against the other top models in the same category.
2. The pricing of the new model against the other top models in the same category.
Grok, Meta, Gemini and others all released updates to their models within around a month or two from their respective last release and made significant jumps in benchmarks all around the same time. Any guesses as to why that is? Is it just the release season and/or everyone is benchmaxxing?
its essentially the same model being trained continuously 24/7 with the company periodically publishing just a new checkpoint
each new checkpoint can benefit from better reasoning training, RL on specific tasks and more synthetic data
So why do they seem to release around the same time ?
my guess is because they time major releases around quarterly earnings, investor meetings and other important business milestones.
Once one company announces a major update, the others also have an incentive to ship their latest checkpoint rather than look like they r falling behind.
Sure, they are just checkpoints, that much I guess is obvious. The question is why did they not do frequent releases like this before and why are they making significant jumps in benchmarks so fast and all these companies suddenly falling into that pattern? Earning reports are not to come until end of October, that's not it.
possibly the beginning of the recursive feedback as models begin to aid in their own improvement? especially algorithmic improvements, which seems to have a lot of wide open space for gains
I'm really curious about this: the foundational paper behind today's LLMs came from Google, and some of the world's best scientists were at Google. So why are they falling so far behind in the AI race?
The "let's make money by selling/renting out TPUs" faction has won and the "let's make money by training and selling a frontier model" faction has lost.
And it's arguably not crazy, at least if SemiAnalysis's estimates are to be believed:
* 20% of all TPU shipments from Q3 2026 through Q4 2027 are sold to SPVs serving Anthropic ($150B of contracted revenue); vs
* ~$12B ARR for Gemini.
> The "let's make money by selling/renting out TPUs" faction has won and the "let's make money by training and selling a frontier model" faction has lost.
It's a variation of opportunity cost. A company that has an opportunity to take $1 and make $1.50 on it can't justify an opportunity to spend $1 and make $1.25, even though a less profitable company may make a good living on that. When considering capital allocation, Google has to consider the opportunity cost of investing more into their highly lucrative ads business. Another company that has no access to such a lucrative business uses different opportunity cost when it comes to allocating capital. It can easily be the case that Google could end up justify being in the business of renting out shovels and end up chased out of the business of using the shovels to create AIs entirely because that turns out not to be where the money is. I'm not saying that's obviously inevitable; I'm saying it's a possible and reasonable outcome.
That's why even though the industry produces giants, these giants can never just eat everything. Even though it seems like they have all the money, it isn't practical for them to try to do everything and in fact limits get hit very quickly for anything other than the primary, lucrative business.
Apparently there is no snappy term for this in the business space, according to such AI searches as I have run.
With all due respect, did you read my comment beyond the first paragraph? It addresses both points, TPU economics/pivot to sales + internal shortages making it hard to train models, to the extent they can be addressed based on public sources.
There are other factors at play, but they're more recent/second-order.
Why does Google keep announcing these subpar models?
What am I missing?
They can't seem to be able to produce a frontier model, fine.
Just be quiet about it and work hard until you manage to put one together.
[EDIT]: Come to think of it. Maybe they're trying to build the Toyota corolla of AI ... let's see if that wins them the battle long term. I personally doubt it.
Original images: https://image.non.io/neonRamenDesigns.webp
Gemini 3.7 build: https://html.non.io/neonRamenGemini3.7
Opus 5 build for comparison: https://html.non.io/neonRamen
Opus is still best in class for this, but it's worth noting how well Gemini 3.7 does vs a more comparable LLM price wise, which is Grok 4.6: https://html.non.io/neonRamenGrok4.6 . I thought Gemini would blow Grok out of the water (it generally has in the past), but Grok has really caught up.
It's not a bad model by any means, but I just don't know what situation I'd reach for 3.7 Flash first for. Google really needs a differentiator, especially given how hard it is to get an API key from them. They can't be high friction and non-pareto.
Disclaimer: I work in Google so it might be that this link is not publicly well known
At a high level though, as a rule of thumb Google assumes that they're serving companies at Google scale first, and at a human scale second. For other companies it's the opposite. Generally what that means is the first experience you get with a Google product will route you through 8 different dashboards to set up ACLs before you've hired your 2nd employee.
Also:
> Google assumes that they're serving companies at Google scale first
So much this. I'm currently grandfathered in until the end of the year on Google's Search API, but the $35,000 they want to continue usage of my < 1000 personal searches per month, not going to happen. It has honestly been easier to use Anthropic to help me build my own search index & crawling infrastructure than deal with Google.
I think that's actually a very interesting insight that would be helpful for PMs on GCloud to take note of. As a single founder, setting up Google Cloud, it's like they start out by assuming you're bigco, forcing (I assume most) of their users into a arduous process of removing components they don't need.
Google AI Studio is one of Google's solutions to this problem, but in typical Google fashion, it's bolted-on without any clear connection in the ecosystem. If you're also using GCloud, it's hard to remember it's even there.
OpenAI's platform, by contrast, is streamlined, easy to use. With Google, I feel like I need to wade through the documentation first before even using the darn thing.
This is the problem.
> Google AI Studio is one of Google's solutions to this problem
This is also the problem.
Where's Sergey's "founder mode"?
“Failed to create project, The request is suspicious. Please try again” or “ You do not have permission to create a key in this project”. You can then navigate multiple screens in GCP to make it work but it’s a hassle compared to any other provider (OAI/Ant/OpenRouter or any of the Chinese labs).
It was literally two clicks, and didn't even leave the page: the dialog asked to create a project and type in a name, I did that, clicked submit and then it was selected as the default project. One more click and I had the free tier API key.
Not saying you didn't have that experience at the time, but personally I have had zero issues with AI Studio and consider it the most dead simple/fastest dev dashboard to get started compared to the others like OpenAI/Anthropic (thanks to Google's free tier that lets you skip billing setup annoyances just to play around with Gemini).
Despite what HN threads (that are also frequently confused and talking about GCP instead) portray as universal/widespread issues or the process being complex and time consuming somehow.
Eventually I gave up and run a few hundred million tokens (edit a few billion) through openrouter.ai using Gemini Flash 1.5 to Flash 2.5
Every since price increases on Flash 3.0 I've stopped using Gemini, too expensive for basic classification, sentiment detection, ocr etc.
As other posters said Google assumes you are some bigcorp trying to use their products. The Vertex versus AI studio confusion did not help.
Sorry but it's not worth waking up with a 100k$ bill, fix your platform first.
Biggest traffic day of the decade and our site was down because of google.
https://ai.google.dev/gemini-api/docs/billing#spend-caps
Using Google products in general is an effing nightmare as soon as you have to give them money.
The one thing you want in a business is to remove friction when people want to give you money, a concept Google has never been able to understand.
Spending money via Google Pay on Android is extremely easy, Google does know how to accept customer's money (in the consumer space)
Google being Google, their models tend to be better at finding, organizing and presenting information, from my experience.
I'm not sure what event-based traders are doing now, but back in the day NLP sentiment analysis was all the rage, so I'm assuming they've now incorporated LLMs too.
You reach for it every time you do a Google search
Maybe things there have improved some, but when I was looking it was a huge runaround.
What does this mean? Anybody can get an API key
I'm used to incremental Figma wireframe -> final product and working together with a designer.
EDIT: OKAY I see it's mostly the "image" generation, not so much the HTML... Noticeable in the food photos and the foodtruck/cart photo
i do wonder why gpt sol was not compared here but honestly it's not really known to be the best at UI
a fable 5 comparison would've been also interesting and likely the best.
I guess I need to try harder. :)
Opus can't generate images since A\ doesn't have a diffusion model.
This is the build step generated by my diffusion-based ui tool's copy-for-agent action.
The grok test was ran through the cursor cli agent however.
https://image.non.io/12275ee8-71e9-4941-823b-e51fec157b4d.we...
The agent is told to generate assets as part of the buildout. It gets to decide what the prompt is for them / whether to do postprocessing like background removal / what type of asset to generate.
Plus is the ramen in HK even any good?
This is testing if it can match visuals that have already been established, and represent them with all the tools available to a web developer. The ramen example was chosen in particular because there are a lot of things that aren't easy to do with CSS, and require creative strategies: 45deg button cuts, angular repeating pattern elements, blending of raster art and svgs, microglyphs, low contrast subtle elements, etc.
Don't ask yourself whether it's a good design, as yourself whether it's a good test.
It's scheduled to double in price on December 31, 2026, but who would anticipate still using this model five months from now? Especially since 3.6 Flash came out just three weeks ago!
My first effort with default thinking level produced an ambitious pelican, let down by a flawed bicycle: https://tools.simonwillison.net/markdown-svg-renderer#url=ht...
Then I ran it on high, medium and low thinking levels (oddly minimal is no longer an option, which WAS an option for 3.5 and 3.6) and got a pretty excellent pelican for the first two:
https://tools.simonwillison.net/markdown-svg-renderer.html#u...
UPDATE: That was in Safari, but as pointed out in the replies here the pelicans do NOT render well in Firefox or Chrome! Best guess is that's because of this invalid filter in the SVG:
Filters are meant to contain additional elements, not be empty: https://drafts.csswg.org/filter-effects/#FilterElement - so maybe Chrome and Firefox remove the element that references the broken filter but Safari doesn't?This suggests you primarily use Safari.
While the bike renders, the pelican doesn’t in Chrome and Firefox.
Probably one of the more serious defects I’ve seen with the pelican. It’s one thing when animated SVGs have bugs, but another when plain ones do.
Anthropic for instance announced a couple of days ago that they are making Sonnet's 'introductory pricing' permanent https://xcancel.com/claudeai/status/2086891169217122586
Those kind of workloads would be hit by an end of introductory pricing. And it's exactly the kind of cases that are not very price sensitive. Where we are price sensitive we track new model releases closely, where we aren't other issues get priority as long as llm performance is good enough
Or maybe one of your installed extensions modifying the page (presumably private tabs disable those)?
I work for Google and their free, internal Gemini API isn't quite as graceful. They once turned down a model arbitrarily and it broke our tests. I had to scramble to fix it, then build warning systems for the turndown as well as a special validator to make sure any upgrades make the same determinations.
https://deepswe.datacurve.ai
> Starting January 1, 2027, $1.50/1M input tokens and $7.50/1M output tokens will apply.
Compare this to Luna which is at $0.2/1M input ($0.02 cached) and $1.2/1M output.
https://developers.openai.com/api/docs/models/gpt-5.6-luna
However, I've noticed 2 drawbacks with Luna. Context rot is much more palpable than Terra and Sol. It tends to get confused and go into rabbit holes when it's context gets filled up. In addition, when instructions are vague, it performs poorly and tends to write way to more code than necessary, but that is to be expected of smaller models. In all, for clearly defined, bite-sized coding tasks, Luna's price-to-performance has been insane. It might have very well commanded the price tag of Sol if it came out just a year ago.
I'll use it locally too, but we use Cursor for work
I don't think I want a gastown-style "just yolo everything" approach, in fact I really want more control over how decisions are made and with what info. Does this exist?
On complete, it moves the user story from Doing to Done. I also have a MEMORY.md file that the agent read and writes in the beginning of a new conversation and at the end of our conversation to update stale information. These files are referred to every time I start a new conversation.
Regarding forking, I know Codex has such button underneath each message that lets you fork the whole conversation. I usually do that when I want to sidetrack and discuss something.
I use no SKILLS or commands like /goal. I’ve come a long way with just prompts and markdown files. Its all a different way of encapsulating instructions anyways.
How much does that matter if it's reset at every turn?
Edit:
I've realised I was incorrect, the thinking doesn't get passed back and forth but the latent snapshot does which result in using memory just the same.
Gemini[0] for example passes along a snapshot of the reasoning state but it's not the equivalent to keeping all the reasoning tokens in the context.
[0] https://ai.google.dev/gemini-api/docs/thinking#signatures
Edit: Apparently it does take the same space in the LLM latent space so I was wrong.
I've always considered the Flash series of models to be for low-cost, high-volume, mostly text-based use cases (e.g. summarization, parsing, formatting), emphasis on low-cost.
[edit: ah, benchmarks here: https://blog.google/innovation-and-ai/models-and-research/ge...
more of a Terra than Luna competitor which is an interesting positioning. I feel like differentiation at the mid-tier of models is pretty difficult.]
I pay for ChatGPT and Gemini, and while Sol is a total beast with anything text, it still poisoned my cucumber bed. Which I will be bitter about for at least a few years while the bed recovers. Gemini (even flash) is exceptionally talented at viewing photos and telling you what to do/what it is (and telling me I just misidentified the problem with my cucumbers and spraying off the "bugs" actually just spread the bacteria everywhere.)
Luna way cheaper. DeepSeek used to be, but I think it's somewhere on Sol's curve after the price hike.
Damn, Luna on max is as good on DeepSWE as Kimi k3, I think I dismissed this model unjustly.
Kimi K3 is a beast though, just costly.
[0]: https://deepswe.datacurve.ai/
Gemini doesn't have adjustable reasoning effort (at least on the graph) so each of its curves is just one point.
So it's better than 3.6 Flash, at half the price. I've been pretty excited about Gemini models recently, they just feel so fast after spending most of the day at work waiting for Opus 5.
I think its the same price..
https://ai.google.dev/gemini-api/docs/pricing today has 3.6 Flash at $0.75/$3.75 until December 31st 2026, then doubling.
https://web.archive.org/web/20260809105129/https://ai.google... Internet Archive copy of that page from 9th August has 3.6 listed at $1.50/$7.50 with no mention of the price changing.
The selling point for gemini continues to be speed and particularly end-to-end response time.
Worth noting that OpenAI just announced that they got the full GPT 5.6 Sol model running on Cerebras at 750 tokens per second. No announcement of the pricing though...
Good catch! You're right to point that out. My previous marketing copy missed that specific detail. Thank you for bringing it up!
https://youtu.be/7NfyZhV1dKM?t=52
The application isn’t so complicated that you need opus level reasoning or code writing, we need “good enough” data retrieval and processing with natural language queries and the ability to answer follow up questions.
For that Gemini works well for a decent price.
This is what I do too.
Edit: and Sol medium actually has the same AA intelligence score as Gemini 3.7, and has >7x fewer tokens, actually making it faster
Unfortunately, it's often not strong enough for heavy refactoring and long running development loops.
I fail to see the usecase where DS V4 Pro is not enough, but Flash 3.7 is - except multimodal.
Luna is similar, and also 8x cheaper. Source: artificialanalysis
The only benefit I can see is the speed, that looks to be outstanding, probably thanks to their TPUs.
Well, compared to 2 months ago, it's no longer 100x more expensive for similar levels of quality...
If they continue monthly-ish releases by 3.9 - by Halloween - they should be close to the best in terms of what you get for what you pay for.
In 2 months, they've gone from basically the bottom of the pack to at least being somewhat usable and competitive.
OpenAI and Anthropic release in a month, and change things. OpenAI is claiming to be close to an Astra release - but that seems like a Fable type release - where they're just releasing a better more expensive model, not more cost effective models.
The only thing that works at scale is gemini flash.
That's why DS4 already had a huge price hike announcement.
Deepseek as a company can just increase prices for the crazily cheap cache they have, that's their only lever.
I could probably do text only for my workflow (feature development/debugging for web microservices) but sometimes it is easier to just toss a screenshot into the Claude prompt, so that gives it an edge.
If your workflow is 100%, certifiably never ever going to involve an image, then yeah, this isn't going to be huge.
Google AI Mode consistently gets me consistently good results and good speeds. It really changes what "googling" is for me.
> Coding and agentic tasks: Significantly higher quality on real-world software engineering and agentic benchmarks, improving issue resolution and reducing failed agent loops.
> Web development and stronger design parity: Generates higher-fidelity desktop and web application code directly from design mocks, with strong gains in design adherence and in auditing existing codebases against mocks to verify 1:1 design parity.
> Promotional pricing: Gemini 3.7 Flash will be available at an introductory price of $0.75/1M input tokens and $3.75/1M output tokens. We’re also applying this new rate to 3.6 Flash. Introductory pricing expires on December 31, 2026; after, $1.50/1M input tokens and $7.50/1M output tokens will apply.
Still no sign of 3.5 Pro. Will have to test it, low expectations given every other model from the Gemini 3 lineage, but one can hope. Just struggle to understand the promotional pricing being temporary for four months. Given this industry, I'd be hard pressed if 3.7 Flash was still in use by end of year, so why not make it the official pricing?
[0] https://ai.google.dev/gemini-api/docs/latest-model
It was probably to placate some kind of general internal pricing/revenue benchmark that doesn't account for new model releases. Politicians do shit like this incessantly and it reeks of bureaucracy.
"Hey, we are not in a race to the bottom. This is our usual pricing, but this now is a promotion because we know we're coming from behind and need to entice users."
They're drawing a line in the sand on monetisation and signalling that to everyone, while in reality offering it a deep discount (no idea if profitable or not) knowing that this model will probably be obsolete before then.
I honestly wonder how long this calliope can keep playing before it crashes to the ground.
(I have no business relationship to anything mentioned here except as a regular retail customer who went bargain-hunting)
OpenAI has too much money. They’re spending their money in stupid ways.
As a business Google might want to focus on fundamental research 2-3 years from now and not compete on who acquires more money losing customers. Just stay little behind and invest money better.
It can regularly cost more than Fable, take longer, and deliver far far lower quality.
I'm much more interested how this compares to Luna - which on price is terribly - but at least on quality the benchmarks make this look competitive / usable.
If Google continues monthly Flash releases like Sundar said they would, and they continue to have this much of an improvement in cost/quality - then in a few months this could reasonably be very competitive with the best of the best.
It is not there yet, but at least it's super fast, I guess.
Less than half its price.
More than 50% discount.
I'm currently using omp with Kimi K3 as the planner and DeepSeek v4 Flash 0731 as the implementer, or CC + Fable for planning and Opus 4.8 for implementation. For API(not coding), I just use DeepSeek v4 flash 0731 and MiMo.
I'm pretty happy where I am, but I'm wondering if these new models provide some new kind of advantage
Later you can make decision to drop low performing ones.
> Introductory pricing expires on December 31, 2026. Starting January 1, 2027, $1.50/1M input tokens and $7.50/1M output tokens will apply.
1) It makes sense to try 3.7 flash before cancelling.
2) Prompting models to be honest is surprisingly effective in my recent experience. But only if they listen to instructions.
today i wanted to link agy to a google cloud project, for that i had to enable 5 different APIs in google cloud UI, then create a subscription for Gemini Enterprise (whatever that is), then link it to a project, then assign it to a user. and after all that, agy couldn't find the subscription.
the best part: i couldn't cancel the subscription. so i just paid $35 for one month and left it.
Personally whenever I use Gemini I've just been using 3.1 Pro because I've had insane trouble with them getting things incorrect like this. Hopefully they'll fix it soon / they've fixed it with 3.7 Flash.
Also have to compare to the recent Grok 4.6 release, which appears to straight up be better AND cheaper
Hard to understand why anyone would choose 3.7 Flash under these conditions.. is Deepmind still a frontier lab?
Per-token cost isn't a great metric given that some use way more tokens than others.
At this point, I think they're mostly targeting Google One and Workspace subscribers, except doing worse compared to Microsoft because they don't have Microsoft's huge enterprise moat built from their DOS and Windows days.
It did fine on my usual benchmark about configuring old Sparc hardware, maybe output slightly faster than before. Even included something new to check in the firmware.
GPT-5.6 or Claude models haven't delivered to me non-running code in ages.
Whenever I have Gemini in the flow, it's fast, but mistake riddled. I have low confidence in the output.
I've had some success with Opus driving Gemini models. It's pointless for GPT family since Sol is cheap enough or can drive terra/luna for arguably better performance, same speed, and better outcome.
As for all of the talk in this thread about modalities. Every SOTA model takes screenshots and verifies work now. Grok-4.6 does this, Luna does it, etc. They can also all work _from_ a screen shot or mockup provided.
I don't think it's a major selling point when every model can do it well and reasonably fast.
That said, eagerly awaiting "pro" and improvements to antigravity.
I've gotten good results with it, but it definitely is more hands on.
Large lumbering enterprise with massive inertia. Where innovators leave as soon as they get a better offer.
None of the authors of the seminal "Attention Is All You Need" paper are still at Google.
Fast forward a decade and Google will be reduced to hiring the kind of mediocrities who deign to work at IBM and Accenture.
This feels absurd to me (my gut is "I want the SMARTEST model I can get!!"), but often I find that my experience of using a flash/sonnet model for every-day workhorse coding they are better.
Its not the same thing, but when I think of that I am reminded of working with some engineers in the past who are incredibly smart and have PhDs (or to put it another way, over-qualified) and they were crap engineers because they'd just not be able to focus on the task and ONLY the task at hand and would get easily distracted by the "why" or "more interesting" things when I just asked them to fix a simple bug or whatever. Again, its not the same thing at all, but it certainly comes to mind when I think of this or experience a pro/opus model suggesting we make huge refactors when a tactical fix is all that is required etc.
Of course, the opus-sized models are great when it comes to huge comprehension/research/debugging efforts where the deeper reasoning is actually useful.
If it implements something simple like a file export, it just knows that the file should have a meaningful name. Vibecoded feature beats most software's lazy "untitled.png".
So, yes, I want the smartest model even for simple stuff. Maybe especially for simple stuff because the tokens burned will be trivial so the cost doesn't give lower models a comparative advantage.
Consistently, lower intelligence models provide worse results in my own work. But I don't have evals on my side, just vibes.
Google probably crunches more tokens daily than the other labs combined, just because basically the entire global population uses Google (sans china) and Google has shoved Gemini into everything.
I wonder if this counts as evidence against that hypothesis? That multi-modal is struggling to keep up with SotA and the best they can offer is competent and fast?
this is hilarious. it is not 2025 any more, by Jan 2027 there will be at least 3 newer generation of models (from other provider) released already. nobody would use flash 3.7 at that time.
sure we used to cling to gemini models in the past, demanding 2.5 models to continue to serve, but since google betrayed us with those price hike, people already spent their time making their production pipeline less dependent on google since then.
heck, even now I'm not sure I even care if they cut the pricing even lower. there are too many models with cheaper price and similar performance now.
They should call it 'face saving pricing after we realized just how terribly did we mis-price the flash 3.5'
> since google betrayed us with those price hike, people already spent their time making their production pipeline less dependent on google since then.
This is my first hand experience. I spent at least $3000 on gemini-3-flash-preview. And exactly $0 total on (3.5+3.6+3.7)
a decent model with a decent harness will determine when the knowledge base is lacking and attempt to fill the holes; thus the good general models can be very easily brought up to speed on niche domains.
Some model are more aggressive by default, some are more verbose by default. To get the result you want for your specific application, you run experiment with prompts and parameters.
Think like a regulator.
and now we have ai agents to automatic migrate the system with new models. in the past we would need to spend hours to design the prompts, then test the output, then write codes to babysitting it. nowadays any ai agent can do it effortlessly.
Personally, I feel like Google blundered on their pricing because while I was using the free version of the Gemini harness, they took away most of the free limits and made people move over to their Anti-Gravity harness for no apparent reason. I was about to splurge for a Pro sub since I already used Google for extra storage but putting up limits like they did made me not want to trust they wouldn't do more price shenanigans. Now their models are behind and it seems like they're scrambling.
Introductory pricing until December 2026 implies no significant Gemini Flash developments until the next year.
Gemini 4 is apparently just around the corner so unless there's a 3 month delay... there's at least a new Flash update.
But, I think it’s also based on what they are being used for, most LLM users are still mainly SWEs or similar as I understand and there’s a ton of data to train them for coding.
My point is mainly that was never the pitch that got ai the hype it did and imo doesn't justify the valuations even if we all lose our jobs to ai. Because it no longer seems like they even think its making other jobs go away.
They compare it to 5.6 Terra, however https://cognition.com/frontiercode puts Terra at about 1/2 the price
Also have to compare to the recent Grok 4.6 release, which appears to straight up be better AND cheaper
Hard to understand why anyone would choose 3.7 Flash under these conditions.. is Deepmind still a frontier lab?
Also impressed with Grok for some stuff.
Somewhere in the same neighborhood as GPT 5.6 Tera and Sonnet 5, depending on the bench.
And yet the entire thread here is people bitching that it's neither 5.6-sol nor Opus or Fable 5.
BTW why are OpenAI and Anthropic even releasing models like terra/luna and Sonnet?
Why? Just why?
Is there a... market?
For you can't have it both ways: either Sonnet and terra/luna make zero sense for Anthropic and OpenAI or Google is a player.
I hope so. It seems mind boggling to me that an user needs to surf around different sections (plural) of google cloud console, then this Vertex and do a dozen clicks to issue a simple key.
In my evals 3.6 Flash (pre price change) was usually a bit more token efficient than 3 Flash, so I‘m expecting same or even lower cost-per-task on 3.7.
Maybe a play by Google to deprecate 3 Flash soon.
Google's Opus competitor is 3.1 Pro Preview which is essentially obsolete (competed with Opus 4.6). They do not have a Fable/Sol competitor.
I guess Google's betting on consumers being price-elastic (preferring to tradeoff intelligence for significant cost savings)
FWIW neither is xAI, there is no "big 3". xAI has had momentary peaks (I think they are having one right now) but they have never been able to claim to consistently push the frontier in any particular direction. You can also infer they aren't a frontier lab from the fact that they sell their compute.
I also think Google is still the best at fitting the most overall intelligences into their models, but for some reason it seems like the model architecture is just bad.
Theoretically there is some difference between Fable and Opus or Grok and GPT, but at the end of the day I'd look at the bottom left of my screen and to my amusement find out that for the past 3-4 hours I've been using model ______.
If the results are semi-decent, I'd keep it on, if not - I'd randomly switch the model and try again.
Actual thing that would affect my selection would be a number of unused tokens I have left for a model ____ for this week.
Maybe it's cause I'm using those for programming and log parsing and all of them are decent enough, but other than that - there are no leaps I see.
1. How the new model performs against the other top models in the same category.
2. The pricing of the new model against the other top models in the same category.
We still don't have a 3.5 Pro, and along comes 3.7 Flash?!
For Google, this is still gemini-3.1-pro-preview, right?
Flash is better than Pro for now.
Path A: Deprecated, do not dare use
Path B: Beta, do not rely
but luna is hard to beat @ capability / cost
Same training dataset, same software, same hardware, same architecture...
I'm wondering what they changed actually for the model to be more powerful if the benchmark results are real and relevant.
Maybe just tweak settings or the reasoning prompts and called it a new version of their model?
each new checkpoint can benefit from better reasoning training, RL on specific tasks and more synthetic data
So why do they seem to release around the same time ? my guess is because they time major releases around quarterly earnings, investor meetings and other important business milestones. Once one company announces a major update, the others also have an incentive to ship their latest checkpoint rather than look like they r falling behind.
And it's arguably not crazy, at least if SemiAnalysis's estimates are to be believed:
https://newsletter.semianalysis.com/p/gemini-is-cooked-but-g...Because they compete for the same scarce resource, the result is a resource crunch for the group that's lost: https://www.latimes.com/business/story/2026-05-18/inside-ai-...
Citation needed.
also, why can't a massive company do two things?
It's a variation of opportunity cost. A company that has an opportunity to take $1 and make $1.50 on it can't justify an opportunity to spend $1 and make $1.25, even though a less profitable company may make a good living on that. When considering capital allocation, Google has to consider the opportunity cost of investing more into their highly lucrative ads business. Another company that has no access to such a lucrative business uses different opportunity cost when it comes to allocating capital. It can easily be the case that Google could end up justify being in the business of renting out shovels and end up chased out of the business of using the shovels to create AIs entirely because that turns out not to be where the money is. I'm not saying that's obviously inevitable; I'm saying it's a possible and reasonable outcome.
That's why even though the industry produces giants, these giants can never just eat everything. Even though it seems like they have all the money, it isn't practical for them to try to do everything and in fact limits get hit very quickly for anything other than the primary, lucrative business.
Apparently there is no snappy term for this in the business space, according to such AI searches as I have run.
There are other factors at play, but they're more recent/second-order.
What am I missing?
They can't seem to be able to produce a frontier model, fine.
Just be quiet about it and work hard until you manage to put one together.
[EDIT]: Come to think of it. Maybe they're trying to build the Toyota corolla of AI ... let's see if that wins them the battle long term. I personally doubt it.
If so, now I understand why they didn't want to release this model