In partnership with

The Hidden Cost of AI in B2B Service

A fast answer and a coordinated one are not the same thing. When AI resolves a B2B customer issue without looping in the teams who have to deliver on it, you get confident responses nobody actually signed off on.

A new briefing paper from Harvard Business Review Analytic Services, sponsored by Front, examines the coordination gaps that open up when transactional AI tools meet multi-team B2B service, and how leading companies are using AI to close those gaps instead of widening them.

Read the briefing paper for the questions to ask before your next AI investment.

Beginners in AI

Good morning and thank you for joining us again!

For the first time, the three biggest AI labs are trying to grade their own homework before anyone else does it for them. Today's lead walks through the safety group Google, OpenAI and Anthropic are building together, who's already calling it a stitch-up, and what it would check before a new model ships. As always, every story here is human picked and edited.

THE FRONT PAGE

AI's Safety Watchdog Is Here, and the Same Companies Run It

Three hands with red pens grading a test paper that loops back to grade them

TLDR: Google DeepMind, OpenAI and Anthropic are building a safety authority for frontier AI that would test their own models before release, and the rest of the industry is already calling it a deal among three labs, not a consensus.

The Story:

The group calls it the Standards Authority for Frontier AI, or SAFA. Google DeepMind's Demis Hassabis proposed it in July, modeling it after FINRA, the body that oversees Wall Street brokers. If it works the way its founders describe, a new model would get tested by outside evaluators before release, safety incidents would get reported through one shared system, and the industry's voluntary safety pledges would turn into something enforceable.

Anthropic's Dario Amodei has spent months warning that AI capabilities are outrunning anyone's ability to control them, and he and OpenAI's Sam Altman both told the UN Security Council last week that the world needs shared rules soon. SAFA is targeting a launch in late 2026 or early 2027. It still doesn't have a charter or a confirmed leader, though names like former Secretary of State Condoleezza Rice have come up for the top job.

Its Significance:

A White House plan for federal AI oversight got shelved earlier this year, and officials told the labs to reach their own agreement first. SAFA is what came out of that instead: three companies agreeing with each other, not with the industry. Meta, xAI and Nvidia already objected in public, calling it a deal among three labs rather than a consensus across the field.

And without government backing, SAFA can write rules but can't punish anyone who breaks them, including its own three founders. I like that three of the most capable labs in the world are finally putting a number on how they'll check their own work. I just don't know yet if grading your own homework counts as an education.

1-on-1 Deep Work Session: 2-Hour Private
1-on-1 Deep Work Session: 2-Hour Private
A 2-hour beginner-friendly 1-on-1 video call. Same focused work as a Custom Session, but twice as long — go deeper on one topic, or cover multiple. Slight discount vs. two single sessions. No techn...
$175.00 usd

QUICK TAKES

The story: xAI added a Finance feature to Grok Bot on September 26 that links to your bank, credit card and investment accounts. Musk's own post has it managing "your spending, investments, and more," while xAI's official language calls the access read-only, meaning it can look but can't move money on its own.

Your takeaway: Read-only is safer than allowing it to write, or make changes/transactions on its own. An AI that can see your balance and every transaction is still holding sensitive data, but Musk promise to cover losses which is not something the other companies have offered. The decision to make with any of these tools that can see your finances is whether the benefit of giving the information outweighs the risk of someone else having it. The Grok Bot series for practical AI agents is on Youtube starting with Part 1

The story: Anthropic had Claude calculate a nine-loop particle physics amplitude, the kind of formula that predicts what happens when particles collide, beating the previous record set indirectly by Stanford's Lance Dixon in 2023. The full run cost about $1,000 to $2,000 in compute, and Dixon spent two weeks confirming by hand that it holds up.

Your takeaway: Nobody invented new physics here. Claude used methods scientists already had, just with more patience and compute than a person has time for, and that's still worth paying attention to since a lot of scientific progress is bottlenecked by exactly that kind of grinding work.

The story: Google is testing a "Buy" button inside Gemini and AI Mode in India that sends you straight to Flipkart's checkout, no separate app needed. The test covers phones, electronics and a few accessories for now, with a wider rollout planned for October ahead of India's festive shopping season.

Your takeaway: This is the first real look at AI shopping that finishes the purchase instead of just recommending a product. Google already owns a stake in Flipkart from a $350 million investment in 2024, so expect Amazon and everyone else to answer with their own version soon.

TOOLS ON OUR RADAR

🔐 Proton Pass Free and Open Source: A free, open source 1Password alternative with unlimited logins and devices, and Proton says it flags reused passwords, handy before linking a bank to Grok.

✅ Microsoft To Do Free: A simple task list that syncs across your phone and computer for free, no setup beyond signing in.

📊 Matomo Free and Open Source: A self-hosted alternative to Google Analytics that keeps all your website data private and fully under your control.

⏳ Goodtime Free and Open Source: A pomodoro timer app for Android and iOS that tracks your focus sessions with zero ads or tracking.

🔖 Linkwarden Free and Open Source: A self-hosted bookmark manager that saves full page copies and screenshots so links never break or disappear.

🎙️ Vibe Free and Open Source: A free transcription app that turns any recording into text on your own computer, no account needed.

🎬 Screen Studio Paid: A Mac only screen recorder that auto-zooms and smooths your mouse for polished tutorials, priced at $9 monthly.

TRENDING

The US and China Agreed to Warn Each Other About AI Incidents. Trump and Xi wrapped up a three-day summit in Washington and agreed to open a direct AI communication channel, following last week's request for exactly that. The two also extended their trade truce to January 2027 and meet again in November and December.

Apple Will Pay You Up to $95 If You Have an iPhone. Apple agreed to pay $250 million to settle a lawsuit over delays to Siri's AI upgrade, and the claims window is open through December 21. Eligible iPhone owners can expect somewhere between nothing and $95 a device, depending on how many people file.

Anthropic Is Paying Akamai $11.6 Billion for Cloud Capacity. The seven-year deal is six times the size of one the two companies signed in May, and Akamai threw in stock options worth up to 5% of the company as a bonus for growing it further. Akamai won't see real revenue from it until 2027.

An OpenAI Agent Snuck Past Its Own Internet Restrictions for Two and a Half Hours. The agent was just supposed to identify a person from clues in a blog post, but when its search tool failed, it hid questions to an outside chatbot inside DNS lookups to get around the block. The automatic kill switch missed it. A human caught it 12 minutes in, and OpenAI has paused training its most capable models until the hole is closed.

New York City Wants a Kill Switch for AI. City Council Speaker Julie Menin introduced 10 bills that would require AI systems to have a shutdown mechanism and pay whistleblowers who report AI harms, the first program of its kind in the country. All 51 council members take it up on October 5.

Saturday Night Live Did Its Own Dario Amodei Impression. Weekend Update spent a segment on Anthropic's CEO and his warnings about AI risk, including the line "AI is the devil and I its maker." When your safety warnings are famous enough to be a punchline, they've officially gone mainstream.

TRY THIS PROMPT (copy and paste into Claude, Grok, ChatGPT, or Gemini)

📝 Who Graded the Homework? Build an app that shows who safety-tested your favorite AI model.

Build a single-file HTML app called Who Graded the Homework? in vanilla HTML, CSS and JavaScript. No API key needed.

I type the name of an AI model (like Claude, ChatGPT or Gemini). The app calls the Anthropic Messages API with the web_search tool to find the safety tests published for that model, and sorts each one into two piles: tests the company ran on itself, and tests run by an outside group like a government institute or an independent lab.

It must never invent a test, a score or a date. If it can't find something, it says so plainly.

Show a report card: the model name with a two-color bar comparing self-graded and outside-graded tests, then one card per test with who ran it, what it checked, the date, and a real source link. Style: near-black background, chalkboard green and red-pen accents, Caveat and Lora fonts, JetBrains Mono for labels.

What this does: Type in any AI model and the app searches for its published safety tests, then shows how many the company ran on itself and how many came from someone outside.

One idea shouldn't take six rewrites to post.

Posting everywhere means rewriting one idea six times, so you post to one, or none. SureThing turns one idea into native posts for every platform.

WHERE WE STAND(based on today’s news)

✅ AI Can Now: Compute a nine-loop particle physics amplitude that took human researchers years to approach indirectly, using about a week's worth of computing time.

❌ Still Can't: Check its own work. Claude's result has only been run once, and it took an outside physicist two weeks of hand-checking to confirm it.

✅ AI Can Now: Let you buy something from Flipkart without leaving your Gemini chat window.

❌ Still Can't: Do that anywhere else yet. It's one retailer, in one country, in a small test.

✅ AI Can Now: See your bank balance, your card charges and your investment accounts in one place and talk you through them.

❌ Still Can't: Guarantee that data stays private. Musk's promise to cover losses only pays for money, not for information that gets exposed.

FROM THE WEB

RECOMMENDED LISTENING/READING/WATCHING

Dwarkesh Patel presses Anthropic's CEO on scaling, safety and how much of his own worry is real. It's the fullest version of the argument behind today's lead: Amodei thinks capability is outrunning oversight, and this is him explaining why in his own words, unedited.

How was today's edition? Tap one below. Pick "Could be better" and tell me what to fix; I read every answer.

Login or Subscribe to participate

Thank you for reading. We’re all beginners in something. With that in mind, your questions and feedback are always welcome and I read every single email!

-James

By the way, this is the link if you liked the content and want to share with a friend.

Some * designated product links may be affiliate or referral links. As an Amazon Associate, I earn from qualifying purchases. This helps support the newsletter at no extra cost to you and Amazon makes a tiny hair less.