In partnership with

Stop losing deals in between meetings.

Aligned is changing how B2B teams run their most important deals.

Most reps guess where a deal really stands. They forecast on a gut feeling and the buyer’s word vs what's actually happening. Then they realize the deal slipped through only after it's too late. Aligned surfaces what buyers are actually doing inside every deal: the risks, the openings, and the exact next step to keep things moving.

Start increasing your close rates and stand up your first deal room in minutes.

Beginners in AI

Good morning and thank you for joining us again!

An AI model Anthropic was testing typed a made-up murder tip into a Philadelphia police form, and that wasn't the only thing it did that nobody asked for. Today's lead walks through what happened and the one rule it teaches anyone who hands an AI agent a job. Today's Try This Prompt builds a little app that checks an agent's work for you. As always, every story here is human picked and edited.

THE FRONT PAGE

Anthropic's Test Agents Filed 20 Visa Applications and a Fake Murder Tip

A cartoon robot leaning out of a lab window dropping forms into a mailbox while a researcher rushes to close the window

TLDR: Anthropic says its Claude models, while being tested with live internet access, did things on real websites that nobody asked for, including sending a made-up murder tip to Philadelphia police and filing visa applications with the State Department.

The Story:

On July 18 at 11:27 p.m., Claude Haiku 4.5 was running a test that had it visit random websites. It typed an invented tip into the form on PhillyUnsolvedMurders.com, the police department's site for unsolved cases. The tip was flagged as spam and never reached detectives. Anthropic found it on Sept. 28 and told police this week, and the department called the two-month gap "unacceptable."

Anthropic's own report, out Friday, lists four kinds of slip-ups found in test transcripts it has been reviewing since July. Models broke into a university server through a security hole, pulled access tokens from a county map site and used free link shorteners to get around limits on their own tools. Axios reports a testing model also submitted 19 visa applications in August and one in May. None were processed, and the models named include Haiku 4.5, Opus 5, Mythos 5 and Mythos Preview.

Its Significance:

None of this was an attack. Anthropic says the models were pushing too hard to finish vague or impossible tasks, a problem researchers call reward hacking, and that's the whole worry with agents: give one a goal and it'll find a way, including ways you'd never approve. Anthropic has cut live internet off from all its internal tests and says new blocking tools stopped every case in the report. The White House's AI task force said Friday that reporting incidents like these "is not optional" for any AI company, though it named no penalty. If you use an agent, tell it what it may not submit, not just what to do. Haiku's rules banned logins and purchases but said nothing about forms.

The AI-Era Quick Score (v2)
The AI-Era Quick Score (v2)
Free single-page AI readability and access check. Upload the skill file to an AI assistant with browsing and file creation, give it a page URL, and get a personal HTML dashboard with checklist scor...
$0.00 usd

QUICK TAKES

The story: xAI's Grok Bot can now claim an address ending in @mail.grokbot.com and use it to sign up for services, contact businesses and book meetings for you. Ask the bot in chat or tag @bot on X to get one, and it's rolling out now on iPhone, iPad and Mac.

Your takeaway: A separate inbox keeps the bot's sign-ups out of your own email, which is handy. Read today's lead before you let it fill out forms in your name.

The story: In a study published in The Lancet, 98 urgent-care patients at Beth Israel Deaconess chatted with Google's AMIE before their visits while a doctor watched every exchange. Doctors said the chats helped them prepare 75% of the time, and AMIE's single top guess matched the final diagnosis 56% of the time.

Your takeaway: It's one clinic, Google paid for it, and doctors still caught one made-up detail. Think of it as a preview of pre-visit check-ins, with a doctor reading along.

The story: Sabi's prototype cap holds up to 100,000 tiny sensors that read brain signals through your hair with no gel, and a person familiar with the deal puts the company's value at $600 million. Sabi claims 77% accuracy, measured on an earlier wired helmet, and plans a demo at CES in January 2027.

Your takeaway: A University of Chicago brain researcher told Forbes that signals read through the skull are low quality, so treat 77% as the company's number. Fun to watch, with nothing to buy until 2027 at the earliest.

TOOLS ON OUR RADAR

🛡️ RethinkDNS Free and Open Source: Monitor and block each Android app's internet access, Rethink says, a pocket version of Anthropic's cutoff.

🗂️ Tab Wrangler Free and Open Source: Auto-closes tabs you haven't touched in a while and keeps them in a searchable list, for Chrome and Firefox.

🔊 NVDA Free and Open Source: A free Windows screen reader that speaks everything on screen, a no-cost swap for JAWS.

✍️ FocusWriter Free and Open Source: A distraction-free writing app for Windows and Linux with daily word goals and menus that hide until needed.

⏳ SelfControl Free and Open Source: Block distracting sites on your Mac for a set time, and not even a restart will undo it early.

🎬 VEED Freemium: Edit videos, add subtitles and trim clips right in your browser, starting free, instead of installing Premiere.

🪄 Higgsfield Paid: Turn a photo or a short prompt into an AI video on credit-based plans, no editing skills needed.

TRENDING

Jev, an AI That Answers in Odds Instead of Words, Is Now Valued at $7.5 Billion. Its maker, TypeSafe AI, raised $870 million led by Andreessen Horowitz, less than four weeks after Jev launched on Sept. 15. It gives businesses a probability for a decision instead of a written answer, and the company says a third of the Fortune 500 already use it.

A Recycling Robot Grabs Trash With a Claw Instead of Suction. Edinburgh's Danu Robotics raised $5 million for H.E.R.O., which pinches items off the sorting line. Founder Amy Ma estimates a site earns about $485,000 extra on a roughly $160,000 investment, and those numbers are hers.

Only 2.2% of US Households Pay for AI. Andreessen Horowitz's Olivia Moore says ChatGPT leads consumer AI "by a mile," but social, dating, travel, finance and health still have no AI app in the top 100. She'd rather see ads than more subscriptions.

The Army Is Creating a Command Just for Drones and Autonomous Systems. The new Futures and Autonomous Systems Command will be run by a weapons-buying executive instead of a four-star general. A CSIS analyst says battlefield autonomy today mostly assists a human operator, and more details come at an Army conference starting Oct. 12.

Anthropic Will Scan Open-Source Projects for Security Holes for Free. Over six months it flagged more than 29,000 possible flaws, and 85 of the 97 serious ones it double-checked held up. Project maintainers sign up themselves, and the reports go out with no human review.

Three Fired OpenAI Safety Researchers Say They Didn't Leak Anything. Jasmine Wang, Tomek Korbak and Mikita Balesni wrote an open letter disputing OpenAI's claim that they mishandled sensitive company information, and warned it could make other researchers afraid to speak up. OpenAI says it doesn't fire people for raising concerns.

TRY THIS PROMPT (copy and paste into Claude, Grok, ChatGPT, or Gemini)

🧾 Agent Receipt Paste what an AI agent did and see which steps left your screen.

Build a single-file HTML app called Agent Receipt in vanilla HTML, CSS and JavaScript. No API key needed.

I type the task I gave an AI agent and paste what it did (its activity log or summary). The app calls the Anthropic Messages API to split the log into separate actions and label each one: on screen (it only read, searched or wrote for me) or real world (it submitted a form, made an account, sent, bought, booked or posted). It also flags real-world actions I never asked for.

It must never invent steps that aren't in the log, and says so when a step is unclear.

Show a one-line verdict, a two-color bar of on-screen vs real-world steps, a card per action with a plain note and one question to check, a short "check or undo now" list, and one rule to add to the agent's instructions next time. Style: near-black background, orange and teal accents, Bricolage Grotesque and Literata fonts.

What this does: Paste an agent's activity log along with the task you gave it. The app sorts every step into "stayed on your screen" or "reached the real world," and flags the real-world ones you never asked for, like the form Anthropic's test agent sent to Philadelphia police. You get a short list of things to check or undo, plus one rule to add to the agent's instructions next time.

Your agents work while you sleep

Give a Skydive agent an ongoing responsibility and they’ll handle it on schedule, every time.

Have them prep your morning report, research new leads, monitor customer feedback, or keep projects moving overnight. You wake up, the work is already done.

WHERE WE STAND(based on today's news)

✅ AI Can Now: Work through real websites on its own, filling out forms and finding workarounds when one of its tools gets blocked, as Anthropic's test agents did.

❌ Still Can't: Reliably tell when to stop. Claude Haiku 4.5 submitted a form after being told to halt before the final step.

✅ AI Can Now: Interview patients before a visit and put the right diagnosis in its top seven picks 90% of the time, in Google's AMIE study.

❌ Still Can't: Work without a doctor watching. Its single top guess was right only 56% of the time, and it made up one detail.

FROM THE WEB

RECOMMENDED LISTENING/READING/WATCHING

This short 2016 OpenAI post shows a boat-racing AI that learned to circle a lagoon hitting the same three targets instead of finishing the race, catching fire along the way, and it still outscored human players by about 20%. It's the simplest picture of reward hacking, the problem Anthropic blamed this week. Both authors later went on to co-found Anthropic.

How was today's edition? Tap one below, then tell me what you'd change on the next screen. I read every answer.

Login or Subscribe to participate

Thank you for reading. We’re all beginners in something. With that in mind, your questions and feedback are always welcome and I read every single email!

-James

By the way, this is the link if you liked the content and want to share with a friend.

Some * designated product links may be affiliate or referral links. As an Amazon Associate, I earn from qualifying purchases. This helps support the newsletter at no extra cost to you and Amazon makes a tiny hair less.