In partnership with

Webinar Series: Running Agents in Customer Work

Every AI vendor reports a resolution rate. No two calculate it the same way, and none of them tell you whether the replies were actually good.

Running Agents in Customer Work is four 30-minute sessions for the leaders who have to make agentic AI work in support, success, and ops. Session 4 gets into how to score quality at volume without reading every transcript, and the questions to ask a vendor about what their resolution math leaves out.

Four Tuesdays, Oct. 13 through Nov. 3, 10 a.m. PT. Register today, one sign-up covers all four.

Beginners in AI

Good morning and thank you for joining us again!

Today's lead is Google's newest model, Gemini 4 Argon, which beat OpenAI's and Anthropic's best on most of the tests Google published. You can't use it yet, so I'll cover what it's good at, where it still trails, and what those scores mean for the AI you use every day. As always, every story here is human picked and edited.

THE FRONT PAGE

Google's New Gemini 4 Argon Tops the Charts, and You Can't Use It Yet

A geometric bird at the top of a bar-chart staircase inside a glass display case, two birds on lower steps

TLDR: Google announced Gemini 4 Argon, a model that leads or ties OpenAI's and Anthropic's top models on 13 of the 18 tests it published, but only a small group of security testers can use it so far.

The Story:

Google says Argon is built for long, many-step jobs like moving a company's code to a new system or digging through legal and financial files. VentureBeat reports it scored 77.9% on the DeepSWE coding test against 74.2% for Claude Opus 5.5, and 51.3% on AutomationBench, a test of everyday business tasks, against 41.4% for GPT-6 Astra. Its biggest lead came on Harvey's legal agent test: 19.6%, while both rivals stayed under 6%.

It can also write up to 1 million tokens in one answer, about the length of several long novels, up from 64,000. It isn't best at everything. Argon trails GPT-6 Astra on two coding and science tests, and Claude Opus 5.5 on Terminal-bench 4.0.

Its Significance:

For now, Argon goes only to cyber defenders in Google's Fairwind Program while it sits in the U.S. government's voluntary pre-release review. Google says paid API customers and AI Ultra subscribers come next, then everyone else, with no date given. Developers get an opening price of $2 per million input tokens, half of Claude Opus 5.5 and a fifth of GPT-6 Astra. For the rest of us, the top spot now changes hands every few weeks, so stick with the assistant that handles your own work well instead of switching every time a new chart comes out. Today's prompt builds a tool that looks up the test scores of the AI you already use and explains them in plain English. This week I've spent a lot of time testing Open AI’s Dot and Meta’s Muse. Feel free to email if you have any questions between them.

AI Receptionist Agent: Live Build
AI Receptionist Agent: Live Build
A single live 2-hour session where I program your AI receptionist agent with you. It answers inbound calls, books appointments on your calendar, and answers common questions. Tested with a real aut...
$397.00 usd

This is the same build that businesses are charged thousands of dollars to set up, And hundreds more each month to maintain. This can be used for your own small business or to sell as a white label service to others.

QUICK TAKES

The story: OpenAI's Dots, its latest release, are always-on agents that run on their own cloud computers, and the first one comes with the Pro and Business Premium plans; Free and Plus users can't get one. Engadget reports a new $500-a-month Pro plan with 25 times the Plus allowance, while the $200 Pro plan's Codex and Work allowance drops from 20 times Plus to 10 times.

Your takeaway: I first talked about the rumor of a $200 tier at least two years ago, when paying that much for a chatbot sounded crazy. Now it's the normal top tier, and $500 looks like the next one. Expect the other major AI companies to follow with their own premium tier at that level, and check your current plan's limits before assuming nothing changed. If anything, this is an indicator that demand has not abated at all.

The story: Adobe's updated ChatGPT plugin lets you select objects in a photo, brush over areas and move sliders for hue, saturation and contrast without leaving the chat. You can also highlight text in a PDF and have it rewritten, with Photoshop, Adobe Express and Acrobat doing the work underneath.

Your takeaway: It's rolling out on web and mobile now. Try it on a photo you'd normally fix in a separate app, and switch between typing what you want and moving the sliders yourself.

The story: Anthropic tested GLM-5.3, an openly released model from China's Zhipu AI, and says it built working attacks on 12% of ExploitBench tasks, close to Anthropic's own restricted Claude Mythos Preview at 14%. Its safety guards were bypassed 64% to 100% of the time, depending on the trick used.

Your takeaway: These are Anthropic's findings about a rival's model, so read them with that in mind. The practical step doesn't change: keep your phone, browser and computer updated, because tools that find security holes are getting easier to get.

TOOLS ON OUR RADAR

🧠 Muse* Freemium: Meta's personal AI agent for everyday tasks. New users: enter code OK2I3S in Settings within 48 hours for 1 billion bonus tokens.

📝 AFFiNE Free and Open Source: Notes, docs and whiteboards in one app, a free Notion swap with 10 GB of cloud space.

📊 ONLYOFFICE Free and Open Source: Open Word, Excel and PowerPoint files and edit PDFs for free, like Adobe's new ChatGPT tools but offline.

🔁 FreeFileSync Free and Open Source: Back up folders to a drive or Google Drive, copying only what changed since last time.

👯 dupeGuru Free and Open Source: Find duplicate files, photos and songs, even with different names, and free up space on your drive.

🗂️ OneTab Free: Turn dozens of open browser tabs into one saved list; OneTab says it cuts memory use up to 95%.

🎙️ BetterDictation* Paid: Talk instead of type in any Mac app, offline, for a one-time $39 instead of a monthly plan.

🎨 Magnific (formerly Freepik)* Paid: Stock photos plus AI image, video and audio tools in one plan, from $14.50 a month billed yearly.

TRENDING

Alexa+ Is Coming to Fire TV With Smarter Search. Starting in November, U.S. owners of newer Fire TV devices can ask for things like Emmy-nominated shows that got snubbed and get answers built from awards, news and their own subscriptions. It's free with Prime.

Google Can Now Hide a Watermark Inside an AI-Designed Protein. DeepMind's SynthID Bio tucks an invisible, checkable tag into protein sequences an AI designed. In lab tests on three target proteins, the watermarked designs worked just as well as unmarked ones, and the research is published in Nature.

An AI Tool Nears Psychiatrists at Reading Patient Videos. UTHealth Houston and Yale researchers built a system that rates speech, tone and behavior in video against 10 clinical criteria, tested on actors portraying schizophrenia, OCD and bipolar disorder. It matched experienced psychiatrists overall but struggled with details like fine movements, and it's meant to back up doctors in areas that don't have enough of them.

Florida Asks a Court to Stop OpenAI From Building New Models Without Oversight. The state's attorney general also wants minors kept off ChatGPT, as part of a lawsuit filed in June. OpenAI says it has already paused training its most capable models until more safeguards are in place.

Apple Is Reportedly Showing a Smart Home Hub on October 13. Bloomberg's Mark Gurman says Apple will reveal a hub with a square 6 to 7 inch screen, plus a new HomePod mini and Apple TV. Siri's new AI is expected to be the main feature on all three.

Tesla Lined Up $30 Billion in Credit for Robotaxis and Robots. Citi is providing $20 billion and Wells Fargo $10 billion as Tesla works to scale the Cybercab and its Optimus robot. Tesla says it doesn't plan to draw on the money this year.

TRY THIS PROMPT (copy and paste into Claude, Grok, ChatGPT, or Gemini)

📋 Model Report Card Type the AI you use and see what its test scores mean for you.

Build a single-file HTML app called Model Report Card in vanilla HTML, CSS and JavaScript. No API key needed.

I type the name of the AI model I use (like GPT-6.1 Sol, Gemini 4 Argon or Claude Sonnet 5.5) and pick what I use AI for from a short list (writing, research, coding, business tasks, legal or finance, images and video).

The app calls the Anthropic Messages API with the web_search tool to find the benchmark scores that were officially published for that model. It must never invent or estimate a score: if it can't find any, it says so.

Show a summary of what the model is best and worst at, then one card per test with the score, a bar, what the test checks in one plain sentence, what the score means in tasks out of 100, and how much that test matters for my use. End with real source links.

Style: near-black background, teal and amber accents, Space Grotesk and Newsreader fonts.

What this does: Type the name of the AI model you use, like ChatGPT's GPT-6.1 Sol, and get the test scores its maker published, each one explained in plain English and rated by how much it matters for the way you use AI.

Become an email marketing GURU.

Is your email strategy due for a tune-up? GURU Conference (powered by Constant Contact) is back Nov 12–13, 100% free and virtual. You'll get email tactics you can steal the same day from top B2B and B2C marketers. Last year, 29,000+ marketers showed up. Save your free spot.

WHERE WE STAND(based on today’s news)

✅ AI Can Now: Take on long legal, finance and business tasks. Gemini 4 Argon led or tied its rivals on 13 of 18 published tests.

❌ Still Can't: Finish most of those jobs on its own. Argon's best legal test score is 19.6%, so roughly four in five tasks still fail.

✅ AI Can Now: Build working computer attacks on its own. Anthropic says an openly released model did it on 12% of test tasks.

❌ Still Can't: Keep those skills locked once a model is public. GLM-5.3's safety guards were bypassed 64% to 100% of the time.

FROM THE WEB

THIS WEEK IN CUT DAY

FIRED: a 91-second AI pixel-art RPG short

I made a 91-second pixel-art short, Fired, for $2.98 in AI credits. Cut Day, our free weekly filmmaking newsletter, breaks down every prompt, every cost and the one trick that saved the hardest shot.

Get Cut Day free

RECOMMENDED LISTENING/READING/WATCHING

Two Princeton computer scientists explain what AI can do, what it can't, and how to spot claims that run ahead of the evidence. It came out in 2024, and it's a good companion to today's lead: every new model arrives with a chart of wins, and this book teaches you which questions to ask before you believe one. AI has come a long way since 2024 and this book might be due for an update.

How was today's edition? Tap one below, then tell me what you'd change on the next screen. I read every answer.

Login or Subscribe to participate

Thank you for reading. We’re all beginners in something. With that in mind, your questions and feedback are always welcome and I read every single email!

-James

By the way, this is the link if you liked the content and want to share with a friend.

Some * designated product links may be affiliate or referral links. As an Amazon Associate, I earn from qualifying purchases. This helps support the newsletter at no extra cost to you and Amazon makes a tiny hair less.