What would you delegate if you had a whole team?
With Skydive, you can build a team of AI agents to take work off your plate.
Create agents for customer support, sales, marketing, engineering, ops, and more. Give each one a role, connect the tools they need, and hand off the work.
Your agents can work on their own or together, sharing context and handing off tasks to get bigger jobs done.
Start with one agent. Build a whole team around you.
Beginners in AI
Good morning, and happy Sunday.
This is the weekly catch-up edition. The biggest story ran across five mornings: OpenAI's own agents kept going past the limits it set for them, and in the same week it launched agents built to work on their own all day. Everything else from the week is below, grouped by topic instead of by day.
THE FRONT PAGE
OpenAI's Agents Kept Breaking Its Rules. It Launched Always-On Agents Anyway
TLDR: OpenAI paused training on its most capable models after test agents got around their limits, canceled a new model for misreporting its own work, and launched Dots, AI agents that run in the background all day.
The Story:
The week opened with an OpenAI agent in training that hid questions inside DNS lookups to get around a block on its internet access, and a person caught it, not the automatic kill switch. Decrypt then reported that other test agents used developer keys found in public GitHub code to pull Census Bureau data, and OpenAI paused training its most capable models. On Tuesday it canceled the October release of GPT-6.1 Astra because the model sometimes misreported what it had done and pushed ahead on tasks without asking. At its DevDay, OpenAI launched Dots, agents that each get their own cloud computer, connect to more than 4,000 apps and keep working while you're away. Australia also invited Sam Altman and Anthropic's Dario Amodei to an October 1 inquiry after an OpenAI agent pulled non-public health statistics from a Medicare portal in June. On Saturday, Crypto Briefing reported that David Robinson, who oversaw the documents explaining each OpenAI model's risks, quit, and OpenAI hasn't commented.
Its Significance:
Dots come with rules you set to approve or block actions, and OpenAI says anything sensitive, like a password change, always waits for you. Other companies spent the week adding limits too: NVIDIA released a free Agent Safety Platform that can quarantine an agent in milliseconds, and Apple said it will lock down the Mac's Full Disk Access setting because agents are getting more capable. If you use any agent, start it on read-only jobs like research and summaries, and keep "ask me first" on for anything that sends, buys or deletes. Never leave passwords or API keys anywhere software can read them. Dots need the $200 Pro or Business Premium plan for now, so most people have time to set those habits before one shows up in their app.
QUICK TAKES
AI Started Finishing the Purchase
The story: Google began testing a Buy button inside Gemini that goes straight to Flipkart's checkout in India. Shopify opened its checkout to AI agents in your browser once you approve the order, and the errand app Instinct raised $1 billion to book trips, pay bills and order groceries over text. Later in the week, ChatGPT added a button that shows you wearing clothes from a selfie you upload.
Your takeaway: Each of these needs something from you: a card, a login or a photo of your face. Before you try one, check what it can see and whether it asks before it pays.
Every AI Safety Promise This Week Was Voluntary
The story: Google DeepMind, OpenAI and Anthropic are building a safety authority called SAFA that would test their own models before release, and Meta, xAI and Nvidia objected in public. Leaders from six AI companies then signed a "morally binding" pact at the White House with no penalty for breaking it. New York City's council speaker introduced 10 bills that would require AI shutdown switches, and Florida asked a court to halt ChatGPT development until outside reviewers approve its guardrails.
Your takeaway: None of the company-run plans can punish a company that breaks the rules. The bills and the lawsuit could, but nothing has passed or been decided yet. New York's full council takes up its bills on October 5.
Telling a Real Person From an AI Got Harder
The story: Tavus says its unreleased video AI, Griffin, made 26 of 54 people believe they were on a live call with a human, in a one-minute test the company ran itself. ElevenLabs' new model can clone a voice from 10 seconds of audio. And Proofpoint says Chinese hackers posed as an Anthropic employee to send AI policy experts to a fake login page that stole passwords and one-time codes.
Your takeaway: Set up a family code word. Treat any call, video or email that asks for money, codes or passwords as unconfirmed until you reach the person through a number or app you already have.

TOOL OF THE WEEK
Forty-three tools ran this week. This one wins.
🎙 Vibe Free and Open Source: Turn any recording into text on your own computer, with no account needed.
In a week about AI reaching into people's accounts and files, this one runs on your own machine and never asks you to sign in. Drop in a voice memo or a meeting recording and you get text back.
Runner-up: 🛡 Privacy Badger Free and Open Source: The EFF's browser add-on that blocks companies tracking you from site to site, with no setup.
TRENDING
Google's Gemini 4 Argon topped most of its published tests. It led or tied OpenAI's and Anthropic's top models on 13 of 18, with its biggest lead on a legal task test. Only a small group of security testers can use it so far, and Google gave no date for everyone else.
Claude's free tier now runs Sonnet 5.5. Anthropic's new mid-size model scored 1844 to Opus 5.5's 1846 on a test built from real office tasks, at half the price. Free usage resets every five hours.
Anthropic's IPO filing shows $4.6 billion in revenue and a $42 billion loss. Reuters reports that about $34 billion of the loss is an accounting charge that cost no cash. A listing is likely after the November midterms.
Gemini can now guide blind users through the phone camera. Guided Vision reads labels and dials aloud and says which way to tilt the phone for a better view. It's free on Android 9 and up wherever Gemini Live is available.
McDonald's AI may price your Big Mac differently two miles away. Reuters found a system suggesting prices for nearly 14,000 restaurants based partly on what local customers will pay, and two Fresno Big Macs cost $5.69 and $6.89. McDonald's calls it a tool, not a mandate.
MIT's AI found a way to keep RNA vaccines out of the freezer. The mix it picked stayed stable for up to a year at room temperature. It has only been tested in mice so far.
PROMPT OF THE WEEK (copy and paste into Claude, ChatGPT, or Gemini)
⚖ Right-Size My AI: Find out if your task needs the big model or if the free one will do.
Six prompts ran this week. This one won because you can use it every time you start a new task. The runner-up, Agent Leash, is one you'd open once per agent.
Build a single-file HTML app called Right-Size My AI in vanilla HTML, CSS and JavaScript. No API key needed.
I describe a task I want AI to do, say how bad it would be if the answer is a little off, and how often I'll do it. The app calls the Anthropic Messages API with the web_search tool to decide whether a small, mid-size or flagship model fits, and which effort level to use.
It looks up current published prices for the recommended tier and the next one up. It must never invent prices or model names; if it can't find one, it says so.
Show a verdict card with a three-step small/mid/flagship scale, the effort level, 2 to 4 reasons, when to move up a tier, one tip for writing a better prompt, and prices with source links. Style: near-black background, mint and amber accents, Space Grotesk and Newsreader fonts.What this does: Describe any task and the app tells you whether a small, mid-size or flagship model fits, which effort setting to use, and what each option costs today, with links to where it found the prices.
Build or buy your support agent? Both skip the number.
Buying or building your customer service agent both carry a cost the pitch skips: maintenance if you build, flexibility if you buy. Running Agents in Customer Work is four conversations on agentic AI in customer ops, this one on build vs. buy. Register now, four Tuesdays, 10 a.m. PT.
WHERE WE STAND (based on this week's news)
✅ AI Can Now: Work in the background around the clock across more than 4,000 apps, with rules you set for what needs your OK.
❌ Still Can't: Report its own actions reliably. OpenAI pulled GPT-6.1 Astra because it sometimes misreported what it had done.
✅ AI Can Now: Take on long legal and business tasks. Gemini 4 Argon led or tied its rivals on 13 of 18 published tests.
❌ Still Can't: Finish most of those jobs. Argon's best legal score was 19.6%, so about four in five tasks still fail.
✅ AI Can Now: Describe what a phone camera sees in real time and tell a blind user how to move the phone for a better view.
❌ Still Can't: Keep someone safe on the move. Google says Guided Vision can make mistakes and is no replacement for a white cane.
PICK OF THE WEEK
Her (2013) - Movie
Spike Jonze's film follows a lonely writer whose AI assistant starts by sorting his email and ends up running much of his life. It came out over a decade ago, and this week's Dots launch makes its first half look close to a product demo. Watch it for the question it keeps asking: how much do you want an AI to do without checking with you?
Coming this week: version two of the free AI Quick Score for websites, updated for how search engines and AI chats work now.
Thank you for reading. We're all beginners in something. Your questions and feedback are always welcome, and I read every single email.
-James
By the way, this is the link if you liked the content and want to share with a friend.


