Your site is losing leads. You can't see where.
Your site looks fine to you — and it's still losing leads. SureThing renders any URL in a real browser, scores 9 dimensions, and ranks the fixes by wasted impact.
Beginners in AI
Good morning, and happy Sunday.
This is the weekly catch-up edition. Everything worth knowing from this past week is below, grouped by what it's about instead of what day it ran. If you missed a few mornings, start here.
THE FRONT PAGE
The Week a Confident Wrong Answer Got Expensive
TLDR: A military analyst's chatbot invented a Chinese ship's cargo and planes were in the air before anyone caught it, a woman's lawsuit says a face match put her in jail for six months, and three separate labs shipped work the same week aimed at getting AI to say how sure it is.
The Story:
On Saturday we covered the intelligence report a chatbot wrote. An analyst at a US special operations command asked it to review reporting on what a China-flagged ship was carrying, it blended public information with secret signals intelligence, and it decided the ship held material tied to nuclear weapons. The analyst used AI again to turn that into a standard report, military planes went up, and armed service members were getting ready to board before officials rechecked it and found the cargo was invented. Friday brought a different version of the same failure: Angela Lipps, 50, is suing for $10 million after facial recognition matched her social media photo to a fake ID, and she sat in jail from July until Christmas Eve, when bank records showed she was a thousand miles away. Her complaint says she didn't resemble the suspect in build, features or tattoos. In between those two stories, three attempts at the underlying problem landed: a former OpenAI researcher released Jev, a model that cannot write sentences and instead picks from a list of options you provide to it and returns a confidence number; OpenAI published six real cases of its own models misbehaving, including one that couldn't find earnings figures for a California county and made them up; and Google DeepMind ran four models through a Nature Machine Intelligence study testing whether they skip a question when their own internal confidence is low.
Its Significance:
The ship report was dangerous because it read like every other report. Nothing in the format tells you which paragraphs came from evidence and which came from the model filling a gap. Facial recognition has the same shape: it returns a match whether or not the match is a real person, and the number attached to it means resemblance, not identity. What's different about this week is that three separate labs put out work aimed at that specific gap rather than at making answers longer or faster. None of them solves it yet, and Jev's own results compare its answers to other AI models rather than to known correct ones. Until something does, the rule holds: open the source before you act on anything an AI hands you, and raise your standard as the stakes go up.

QUICK TAKES
Your Assistant Moved Into Your Personal Files
The story: Apple switched on the rebuilt Siri on September 14, and it digs through your messages, emails and photos to answer one question, and reads your screen on request. The models behind it were built with Google. The next day Google released Gemini 3.8 Live and Live Extended Thinking, which read your camera as you go, detect and switch between 97 languages on their own, and keep talking while a job finishes in the background, shipped straight into Gmail, Docs, Keep and Search Live.
Meta is reportedly preparing glasses called Luna with no camera and six microphones, and Google Home opened up to outside agents like ChatGPT and Claude, which can now read your device history and flip your lights.
Your takeaway: All of these only get useful once you hand over the mail, the texts, the photos, or the camera roll. Siri AI is English-only in beta, isn't in the EU or China, caps the cloud features daily, and Apple says wider access costs money later. Google Home's agent access is US-only on the $20 Premium Advanced tier. And dropping the camera from a pair of glasses doesn't drop the six microphones that are always listening.
Medicine Is Where AI Keeps Putting Up Real Numbers
The story: MIT published a system called xvr that trains on thousands of simulated X-rays built from one patient's own CT or MRI, then matches the live X-ray to that 3D picture with sub-millimeter accuracy in seconds, tested on records from five hospitals covering kids and adults and beating other AI methods by roughly ten times. England's NHS added the Brainomix e-Stroke system to brain scans, and the gap between a patient arriving and getting treated dropped by a full hour, with the Department of Health saying the number leaving with little or no disability tripled. Google's tally of its AI in science included a study with Imperial College London and the NHS where AI caught 25% of the cancers missed between screenings across mammograms from 175,000 women.
Your takeaway: Same technology as the front page, opposite result, and the difference is the job. One narrow task, one answer a person can check against a known correct one, a measured result against a baseline. When you read an AI claim this week or any week, look for those three things before you believe the number.
Everyone Has a Rule They Want Somebody Else to Follow
The story: China's foreign ministry pushed back on Anthropic CEO Dario Amodei after he argued for keeping US limits on advanced chip sales, with the state-run Global Times calling the essay a Cold War playbook. Cato's Jennifer Huddleston and Block co-founder Jack Dorsey argued the same day that a government-ordered pause would shield the biggest labs from competition, and Mark Zuckerberg posted that each company can handle its own safety work. Microsoft published a draft code of conduct and a Brookings and Fudan University report asked Washington and Beijing to agree that only people, never AI, can order a cyberattack on nuclear command systems.
Your takeaway: Every one of these positions happens to favor the person holding it, which doesn't make any of them wrong, it just means you should notice it. Only one of them is open to you right now: Microsoft's comment window runs six weeks, and the finished version shapes how its models are trained in 2027.
TOOL OF THE WEEK
Twenty tools on the radar this week. This one wins.
🔑 KeePassXC Free and Open Source: Store passwords in a locally encrypted database with AES 256 encryption, works fully offline, no account needed and highly regarded.
The whole week was assistants asking for the keys to your mail, your photos and your house. A password vault that lives on your own disk, never asks you to create an account and never uploads anything is the one part of that stack where you can point at the file and say exactly where it is. The only learning curve is picking one strong master password and not losing it.
Runner-up: 📧 Thunderbird Free and Open Source: Manage multiple email accounts in one client with built-in encryption, spam filtering and RSS, no self-hosting needed.
TRENDING
OpenAI is testing ads inside ChatGPT that you can hold a conversation with. Click a sponsored result and you can open a separate chat paid for by that business, ask whether the table seats six, and never leave the app. OpenAI says the sponsored chat is labeled and kept apart from ChatGPT's own answers, and it wired the ad tools into Shopify and HubSpot the same week.
OpenAI also shipped a version of Astra built for lawyers, with a search index covering US case law, statutes, regulations and court rules across more than 230 million web addresses, plus 26 partner plugins. Healthcare, then financial services, now law. If your field runs on expensive research software and hours billed by people, a version of this is coming for it.
Meta pulled its paid features into one subscription called Meta One, covering Instagram, Facebook, WhatsApp and Meta AI in a single plan with more than 50 features. Prices start at $2.99 a month for one app and run to $499 for the largest business tier. Meta says the free versions aren't changing.
Gemini Notebook added live voice study sessions in nearly 100 languages, plus a phone recorder for lectures and auto-built quizzes and flashcards. US college students can claim a year of Google AI Pro free, normally $19.99 a month. Every answer stays tied to the sources you upload, which is the part worth stealing whether or not you're in school.
Anthropic confirmed it runs a real biology lab, meaning physical biological material rather than simulations, in the Bay Area. Its head of life sciences told Reuters the final test in biology is still real lab work. The company bought a stealth biotech called Coefficient Bio in April and has opened access to its strongest models for vetted bio researchers.
The AI actress did a press tour and it fell apart on camera. Particle6 put its generated performer Tilly Norwood in front of 75 journalists, and in a sit-down with Piers Morgan she couldn't follow a simple question about whether her co-stars were real, then spoke Chinese for more than ten seconds before apologizing for a hiccup. The coverage was the point, which is worth remembering when you read the next AI stunt.
PROMPT OF THE WEEK (copy and paste into Claude, ChatGPT, or Gemini)
🏷️ Name what you're shopping for. Find out when it goes on sale, grounded in real current buying guides, not a guess.
Build a single-file HTML app with vanilla HTML/CSS/JS. The Best Time to Buy Calculator — name a product category, get real seasonal sale pattern research via live web search.
Aesthetic: dark near-black green (#0f140f), kelly green (#5cc26e) primary with a green glow top-left, gold (#e0b458) for the discount stat with a glow bottom-right. Inter for headings/body, JetBrains Mono for labels.
Form: single product-category text input (encourage a type of thing, not a specific model).
Technical: call the Anthropic Messages API WITH tools:[{type:'web_search_20250305', name:'web_search'}] (works automatically), system prompt as a savvy consumer shopping researcher who searches for CURRENT, real published buying guides and retail pattern reporting on when this product category typically goes on sale, grounding timing windows and reasoning in real current sources rather than generic assumptions, honest about discount ranges only when confidently found via search. Return raw JSON: product, typical_discount (general range if confidently found, empty string otherwise), best_windows (2-4: window + why, specific to the category), worst_time (1-2 sentences on peak-price timing to avoid), if_you_need_it_now (1-2 sentences of practical advice: price tracking, open-box, negotiating, older models).
Render: a gradient "typical discount when timed right" hero card with the product name and discount figure. A "best windows to buy" list of green-bordered cards (window + why). A red "avoid buying" card. A gold "need it now?" card with practical fallback advice.What this does: Enter a product category, a mattress, a laptop, patio furniture, and it searches the live web for real current retail reporting on when that category drops in price and why, model-year clearances, seasonal markdowns, big sale events, whatever applies to that specific thing. You get two to four real timing windows with the reasoning behind each, a typical discount range when one is confidently found, a warning on when prices peak, and practical advice for when you can't wait.
WHERE WE STAND (based on this week's news)
✅ AI Can Now: Match a live surgical X-ray to that same patient's own 3D scan with sub-millimeter accuracy in seconds, trained per patient rather than once for everybody.
❌ Still Can't: Combine several intelligence sources and tell you when it doesn't know. The ship report was written in the same format as a correct one.
✅ AI Can Now: Hold back an answer when its own internal confidence falls below a line, and shift that line up or down when you adjust it inside the model.
❌ Still Can't: Report that confidence cleanly. The DeepMind paper found both the spoken confidence and the number scores are rough copies of something richer happening inside.
✅ AI Can Now: Watch your camera feed live and detect and switch between 97 languages without being told which one you're speaking.
❌ Still Can't: Stay in the right one through a single unscripted interview, as the AI actress showed when she dropped into Chinese for ten seconds on camera.
Hire anyone, anywhere — compliant in under 3 days
Found the right hire, but no entity in their country? Remote becomes the legal employer — with contracts, benefits, and tax setup handled, plus direct access to the same in-house team that runs payroll locally.
PICK OF THE WEEK
A short, free-to-read 2003 novel by the founder of HowStuffWorks, following one worker's life as a piece of retail management software called Manna spreads from fast food into nearly every job in the economy. It pairs well with a week whose front page was about software making a call no person checked, and whose most interesting new model is one built to make decisions for other software rather than talk to people. Written before large language models existed, and free.
Thank you for reading. We're all beginners in something. Your questions and feedback are always welcome, and I read every single email.
-James
By the way, this is the link if you liked the content and want to share with a friend.


