Google just shipped a hacking model. Here's the diff.
Google shipped Gemini 3.8 Flash ($0.75 in / $3.75 out, price doubles Jan 1) and Gemini 3.8 Flash Cyber — a vulnerability-hunting model it only sells to 650 vetted "trusted defenders" — and four hours later Meta shipped Muse Spark 1.3 with a $0.10-per-million "contributor" tier where the discount is your session transcript.
Google shipped Gemini 3.8 Flash ($0.75 in / $3.75 out, price doubles Jan 1) and Gemini 3.8 Flash Cyber — a vulnerability-hunting model it only sells to 650 vetted "trusted defenders" — and four hours later Meta shipped Muse Spark 1.3 with a $0.10-per-million "contributor" tier where the discount is your session transcript. We check both against the one benchmark neither company wrote. Verdict: SHIP IT.
What this video covers
- Google ships Gemini 3.8 Flash + Flash Cyber (86.2% CyberGym, Fairwind Program)
- Meta ships Muse Spark 1.3: $1.25/$4.25, or $0.10/$0.20 if Meta may train on it
- Three sites made 215,128 fake "best software" pages; Perplexity cites them
Transcript
0:00 Google shipped its third Flash model in six weeks today, plus a version trained to hack, and four hours later Meta shipped a model that costs ten cents per million tokens if you let Zuckerberg read your code, so the AI price war now has a cybersecurity division. It's Wednesday, and the diff is long. Gemini 3.8 Flash hit eleven hundred points on Hacker News, Muse Spark 1.3 took nearly seven hundred more, and in between, a research shop found three websites that published two hundred fifteen thousand
0:28 fake best-software pages purely so Perplexity would cite them. Also today: a post begging you to hang on to your Firefox got nine hundred eighty-eight points, and Mistral's help page on opting out of training got nearly five hundred, mostly from Europeans discovering the sovereign option trains on your prompts by default. In this video: what Gemini 3.8 Flash actually costs, the hacking model Google will only sell to six hundred fifty partners, Meta's ten-cent tier and what it really charges, and the one benchmark neither company wrote.
0:59 It's Wednesday, September 2nd, and this is The Daily Diff. Gemini 3.8 Flash is seventy-five cents per million tokens in and three seventy-five out, the same as 3.7 Flash from three weeks ago, with a footnote saying the price doubles on January 1st, which is a free trial with extra steps. On the slide-deck benchmarks, which are undefeated, it scores 54.9 percent on HLE-Verified, X says it beats Sol and Opus 5 on Terminal-Bench, and Logan Kilpatrick posted a 73.7 on DeepSWE, the one long-horizon
1:29 coding benchmark that hasn't been memorized yet, allegedly. Simon Willison asked it to make a cool thing in HTML and got a particle simulation in thirteen seconds for 1.8 cents, running at sixty frames per second, because the sixty was hard-coded into the page. Google's own post explains the gains with a sentence I'd frame: 3.8 Flash works harder. It executes extra reasoning steps, calls tools iteratively, and, quote, might use more tokens, which is corporate for the meter runs faster. The independent DeepSWE board agrees on both counts: Flash scores seventy-four
2:02 percent, tied with Opus 5 at a fifth of the cost per task, but it needed a hundred sixty-six steps and a hundred forty-three thousand output tokens to get there, more than twice what Sol spent. Artificial Analysis burned a hundred forty million tokens and a thousand seventy-eight dollars just to grade it. Cheap per token, chatty per task, and the first token arrives after twelve seconds of thinking. Then the interesting half.
2:23 Gemini 3.8 Flash Cyber is the same model with the safety rails loosened for, quote, trusted defenders, and Sundar says it hits 86.2 percent on CyberGym, the benchmark for autonomously finding vulnerabilities. It patches too: 47.2 percent on CWE-Bench against 47.8 for a leading frontier model, the Chrome team says it produced 2.6 times more correct patches than much larger commercial models, and Google's cloud researchers used it to find a critical vulnerability in under two hours, a job that usually takes months, or one motivated teenager.
2:54 Google insists it prioritized fixing over exploitation, which is the polite way of saying the model can absolutely walk through the holes, they just asked it not to. So you can't have it. Access goes through a new Fairwind Program: governments, critical infrastructure, and six hundred fifty vetted partners with mandatory two-factor. Every frontier lab now owns a hacking model it won't sell you, and this week it's a cheap one.
3:16 Four hours later, Meta released Muse Spark 1.3, and Zuck called it frontier performance almost too cheap to meter, which is true, with one asterisk the size of Menlo Park. The normal endpoint is a dollar twenty-five in and four twenty-five out, more than Gemini. The contributor endpoint is ten cents in, twenty cents out, cache reads at a fifth of a cent, up to twenty times cheaper, and the only difference is a column that says used to improve our products.
3:40 So the discount is your session transcript, and for once nobody has to read the license, because Meta wrote the license as a price list. Top comment: it is now completely obvious how much stealing my tokens is worth. Second comment wonders how long until someone extracts an AWS key from a model trained on contributor prompts. On Meta's own scorecard, Spark 1.3 gets 75.4 on DeepSWE, above Sol and Opus 5, and Meta says it uses twenty percent fewer tool calls
4:06 and twenty-five percent fewer tokens than 1.2, the exact opposite bet from Google, so today the two biggest ad companies on earth disagree about whether talking more is good. Meanwhile Trellner Research asked Perplexity for the best software in three hundred eighty categories and read every citation: 59.8 percent came from sites outside the top hundred thousand, Wikipedia was cited three times out of seven and a half thousand, and three sister sites with two hundred fifteen thousand
4:34 generated buying guides title their homepages Facts and Grounding Page, addressed to the crawler, not to you. The AI SEO war has begun, and its first weapon is regex with a marketing budget. That's a lot of price tags for one Wednesday; if you'd rather read this than hear me say it, the diff lands in your inbox every morning — free at the daily diff dot dev, link below. So, today's verdict: ship it. Gemini 3.8 Flash ties Opus 5 on the one board Google didn't write,
5:00 for a fifth of the price. Just read the token bill, and read Meta's column headers. That's today's diff. I'm Niko from Axrisi. Merge responsibly.
Sources
- Introducing Gemini 3.8 Flash and 3.8 Flash Cyberblog.google
- Fairwind Programblog.google
- DeepSWE v1.1 leaderboarddeepswe.datacurve.ai
- Artificial Analysis: Gemini 3.8 Flashartificialanalysis.ai
- Muse Spark 1.3 pricingdeveloper.meta.com
- Introducing Muse Spark 1.3research.meta.ai
- Hang on to Your Firefoxwww.newsonaut.com
- Mistral: opting out of traininghelp.mistral.ai
- HN: Gemini 3.8 Flashnews.ycombinator.com
- HN: Muse Spark 1.3news.ycombinator.com
- HN: 215,128 "best software" pagesnews.ycombinator.com



