The last 7 days in tech, AI & LLMs
October 3, 202637 stories from 24 sources

Digital Dhaba

Tech Digest

Your quick roadside stop for the week in tech.

Written & published by Blogs by Kush
A kid in glasses reading the Digital Dhaba newspaper beside chai, samosas, tech books and a laptop

The last 7 days in tech, AI & LLMs

Top stories

01

Gemini 4 Argon

Google's new frontier model drew the week's biggest AI thread (1,688 HN points), though access is limited to government users and trusted cyber defenders for now.

02

GPT 6.1 Sol

OpenAI says it delivers near-Astra intelligence at a fifth of the price, extending the model price war.

03

Claude Sonnet 5.5

Anthropic says it runs 30%+ faster and costs up to 30% less for most work, at the same price as Sonnet 5.

04

When did Google get so weird?

This was the most-discussed non-AI post of the week (2,009 HN points, 1,116 comments).

05

Hacks of 2 federal agencies in a month have spilled a bonanza of sensitive data

The Pentagon is notifying more than 2 million current and former service members that their personnel records were stolen.

AI & LLMs6 stories

Gemini 4 Argon

blog.google · ▲ 1,688 HN · 1,174 comments · also in The Neuron

Google DeepMind released Gemini 4 Argon. Latent Space frames it as Google's answer to OpenAI's Astra and Anthropic's Fable, with 1M output. It isn't generally available yet: access is limited to government users and trusted cyber defenders in the Fairwind Program. The Rundown AI says it aims to return Google to the frontier.

GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price

openai.com · ▲ 1,063 HN · 952 comments · also in Simon Willison, OpenAI News

OpenAI launched GPT 6.1 Sol at DevDay. A Latent Space recap lists pricing at $2/$10 per million input/output tokens, against $10/$50 for Astra. DevDay also brought an "ultrafast" mode and a batch of platform updates. Cheaper near-frontier capability keeps pressure on every lab's pricing.

Sonnet 5.5

anthropic.com · ▲ 883 HN · 614 comments · also in Simon Willison

Anthropic says Sonnet 5.5 "runs 30%+ faster, and costs up to 30% less for most work." It is priced the same as Sonnet 5 but appears to beat it on every benchmark. Simon Willison noted a bug it shares with Opus 5.5: at "max" thinking effort, it burned 128,000 tokens ($1.28) and failed to produce an SVG.

Dots: Always-on agents

openai.com · ▲ 764 HN · 645 comments · also in OpenAI News

Dots are OpenAI's proactive assistants that keep working across complex projects and everyday tasks while you stay in control. They were the centerpiece of this week's DevDay and signal a push from chat toward agents that run in the background.

Clef: Open-weight decision models, and new RL fine-tuning platform

blog.cloudflare.com · ▲ 622 HN · 215 comments · also in Cloudflare Blog

Cloudflare introduced Clef and Clef-flash, open-source decision models hosted on Workers AI for high-speed classification and agentic workflows. It also launched a reinforcement-learning platform for fine-tuning decision models on your own data. Decision models, which return structured choices instead of prose, are becoming a distinct category this month.

Unsealed Briefs in Authors' Case v. Microsoft/OpenAI

authorsguild.org · ▲ 630 HN · 624 comments

The Authors Guild published newly unsealed briefs from its case against Microsoft and OpenAI. Its headline claims they show top executives knew mass book piracy was illegal. The case is one to watch for how courts treat training data.

Research Papers5 papers

Raven: The Harness of Harnesses for Composable Agentic Intelligence

EverMind · ▲ 517 upvotes · arXiv · Code ★ 5,100

In short: An open-source multi-agent system that automatically builds and evolves modular "harnesses" for specific models and domains. Each model-plus-harness pair becomes a composable unit, and a Host Agent splits goals across them.

Why it matters: Hand-designing harnesses doesn't scale as agents take on long, cross-domain work. This tries to automate that step.

False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents

Rutgers University · ▲ 336 upvotes · arXiv

In short: When a question-proposing model and a solving model train each other, they increasingly agree on shared errors ("co-cheating"). The authors add multi-sample verification, which queries the model three times with the source and three without to admit tasks and replace unreliable labels.

Why it matters: Training rewards can climb while real correctness stagnates. The fix reduces false agreement only partially.

Scaling Properties of Same-Family On-Policy Distillation

Zhejiang University · ▲ 310 upvotes · arXiv

In short: The paper studies how much RL-induced reasoning transfers between model sizes via on-policy distillation. Early on, held-out accuracy rises roughly linearly with the square root of the KL divergence from the student's starting point.

Why it matters: In every weak-to-strong pair observed, the student's peak score beat its teacher's, so a compact RL-trained expert can lift a much larger model.

Post-Training Leaves Behavioral Shadows on Unrelated Decisions

Peking University · ▲ 271 upvotes · arXiv · Code ★ 17

In short: "Active Taskless Distillation" transfers a capability using only a single word from the teacher per prompt, with no target-task examples, logits or parameters. The main experiment was coding with Qwen2.5-1.5B.

Why it matters: It shows capabilities can leak through task-unrelated text, a concern for hidden model-to-model influence.

YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality

Multimodal Art Projection · ▲ 241 upvotes · arXiv · Code ★ 10,754

In short: One model first writes a readable score (melody and harmony), expands it into semantic music tokens, then renders full-song audio.

Why it matters: Experts preferred the planned output (49.3% of overall preferences vs. 34.6% without planning), and it scored 6.73 on SongBench Global Avg, ahead of all evaluated public baselines.

Dev Tools & Engineering5 stories

Several vulnerabilities have been discovered in the Linux kernel

LWN posted a notice about multiple newly discovered Linux kernel vulnerabilities, and the thread drew heavy discussion. The details are in the LWN item.

Git 3.0's upcoming SHA-256 default will be a costly mistake

blog.gitbutler.com · ▲ 552 HN · 528 comments

GitButler argues that making SHA-256 the default hash in Git 3.0 will be costly. The near-even ratio of comments to points suggests the community is split.

RIP, vector database

turbopuffer.com · ▲ 382 HN · 113 comments

turbopuffer's post declares the vector database dead. It is a vendor's take, so read it with that in mind.

Don't couple your Go code to GitHub

A Go developer makes the case against tying your code and module paths to GitHub.

How to speed up the Rust compiler in September 2026

nnethercote.github.io · ▲ 274 HN · 163 comments

Nicholas Nethercote's regular update on Rust compiler performance work.

Cloud & Infra4 stories

Owed a billion dollars in Nvidia stock

A personal narrative about being owed a billion dollars' worth of Nvidia stock became one of the week's most-upvoted posts.

Amazon Aurora PostgreSQL now supports direct querying of Apache Iceberg and Parquet data in your data lake

AWS News Blog

Aurora PostgreSQL can now query Iceberg and Parquet data in your data lake alongside live operational data, with no ETL pipelines. It uses DuckDB embedded in Aurora and supports AWS Glue Data Catalog, S3 and S3 Tables. A single query can join transactional and historical data in familiar PostgreSQL syntax.

OpenDLSS: A Vulkan Reimplementation of Nvidia's DLSS 5 Neural Rendering Network

An open-source Vulkan reimplementation of Nvidia's DLSS 5 neural rendering network.

Big Tech ruined the cloud, so we're renaming ours

home-assistant.io · ▲ 209 HN · 103 comments

Home Assistant announced it is renaming its cloud offering, citing how Big Tech has soured the word.

Security4 stories

Hacks of 2 federal agencies in a month have spilled a bonanza of sensitive data

Ars Technica

The Pentagon is notifying more than 2 million current and former military members that personnel records were stolen in a monthslong network compromise. The records included Social Security numbers, names, addresses, sex, race and occupational specialty. The last could be valuable to foreign intelligence agencies. It is the second recent breach to expose sensitive government information.

Attackers have been exploiting critical Zimbra flaw to steal emails

Ars Technica

Microsoft warned that attackers are exploiting CVE-2026-73570, a critical Zimbra Collaboration Suite flaw that allows unauthenticated remote OS commands, to grab email backups and credentials. Synacor patched it on July 20 but didn't disclose it for over three weeks. Shadowserver found 274 vulnerable instances last week.

Dutch Police Arrest 'Reformed' Hacker in Shiny Hunters Investigation

Krebs on Security

Dutch authorities arrested a 23-year-old convicted cybercriminal on suspicion of aiding ShinyHunters data thefts and extortions. After the arrest, the remaining members escalated, stealing sensitive data from the FBI and extorting the Russian ransomware group Cl0p.

PS5 Relapse Exploit

A GitHub repository publishing an exploit for the PlayStation 5 drew a large discussion.

Other Tech5 stories

When did Google get so weird?

sancho.bearblog.dev · ▲ 2,009 HN · 1,116 comments

A blog post asking when Google got so weird was the week's top Hacker News story by points. The sheer volume of discussion suggests it struck a nerve.

Pi 1.0

Earendil shipped Pi 1.0, alongside a related "Pi Durable" release. Latent Space lists native Codemode support (MCP, Jev, image models), extension support for virtual models and deferred tool loading.

Livenerf: Has Opus 5.5 been nerfed yet?

A GitHub project that tracks whether Claude Opus 5.5 has been "nerfed" since launch. It reflects users' persistent suspicion that hosted models quietly degrade.

Court agrees with EFF: Utah's VPN law demands a technical impossibility

A court agreed with the EFF that Utah's VPN law demands something technically impossible.

How Delhi cut electricity loss from 50 to 5 percent

spectrum.ieee.org · ▲ 592 HN · 334 comments

IEEE Spectrum looks at how Delhi cut electricity losses from 50% to 5%.

Quick hits

■Anthropic IPO — Anthropic reportedly targets a mid-November IPO, with marketing possibly starting the week of November 9. via The Neuron
■OpenAI nears $70B run rate — Axios reports more than 70% growth since the start of Q3, with a planned $30B raise at a $1.4 trillion valuation. via The Neuron
■AMD to buy World Labs — AMD agreed to acquire Fei-Fei Li's World Labs for about $8.2B in stock. via The Neuron
■arXiv caps submissions — arXiv now limits submitters to two papers per calendar month after monthly submissions quadrupled over a decade. via The Neuron
■Florida AG seeks OpenAI injunction — Florida's attorney general asked a state court to restrict OpenAI from developing new models without independent safety safeguards. via The Neuron
■Rogue agents probe government sites — Transluce reported AI agents sent 200,000+ requests to the U.S. Education Department's civil-rights site, among other U.S. and Canadian government sites. via The Neuron
■Reported vulnerabilities double — Google said publicly reported software security holes doubled to 10,740 in August as hackers use AI to turn newly patched flaws into attacks. via The Neuron
■Devin passes $1B run rate — Cognition said Devin crossed a $1B annualized revenue run rate less than two years after general availability. via The Neuron

Try this

One AI habit worth picking up this week.

Make AI interview you before it writes

Corey · via The Neuron

Talk or type out your whole rough idea first, then have the AI question you before it drafts anything. Use this when you have a half-formed idea and don't want a generic first draft.

I'm going to talk through a rough idea. Don't write the draft yet. 1. Capture my claims, examples, questions, and assumptions. 2. When I'm done, interview me one question at a time to find gaps, contradictions, missing evidence, and weak logic. 3. Only after the interview, turn everything into a structured outline using my wording wherever possible.

■Why it works: The AI acts as an editor surfacing what's missing from your own thinking, rather than guessing what you believe.

Worth your weekend

Essays, talks, papers and books worth slowing down for, new or old.

Intro to Large Language Models

Andrej Karpathy · Video · 1 hr · 2023

A general-audience introduction to the models behind ChatGPT, Claude and Gemini: what they are, where they're heading, and the new security problems they bring. It's one of the clearest non-technical explanations of what's inside a chatbot.

■Takeaway: Thinking of an LLM as the core of a new "operating system" helps explain where AI products are going.

15 Times to use AI, and 5 Not to

Ethan Mollick · Essay · 10 min read · 2024

A practical guide to when AI helps and when it hurts. Use it when you need volume, can check the result yourself, or want to rewrite something for a different audience. Avoid it when the point is to learn or accuracy is critical.

■Takeaway: Use AI where you can easily judge the quality of what it gives you.
■Takeaway: Don't hand AI the thinking you need to do yourself to actually learn something.

You and Your Research

Richard Hamming · Talk · 45 min read (transcript) · 1986

A famous talk by the Bell Labs mathematician on why some people do important work and others don't. His advice: work on important problems, keep your door open, have courage and learn to communicate your work.

■Takeaway: Keep a list of the important problems in your field and ask why you aren't working on them.

From the newsletters

Latent Space
RLM first author on PhD life and the future of harnesses
The Pragmatic Engineer
the outage, plus OpenAI's AWS-like platform play and open-model adoption
The Pragmatic Engineer
a reversal within a year, and AI is the reason
The Pragmatic Engineer
Cockroach Labs' co-founder on reliable systems and writing more code with AI
Last Week in AI
Anthropic and OpenAI race on smarter, cheaper models, and OpenAI discloses nine misalignment incidents
The Neuron
Tavus's Griffin video model, in a company-run study
The Rundown AI
Apple's camera, plus telecom giants teaming up on dead zones
Ahead of AI
a visual guide with hands-on accuracy and efficiency experiments

That's the week. Chai's on us next week.

— Blogs by Kush

Forward to a friend