The Daily AI Show: Issue #113

We are 3 . . . and growing!

Welcome to Issue #113

Coming Up:

Google Earth's AI Lasted One Day

They Trained AI to Refuse Consciousness. It Spread.

The Agent Made Promises. The Prompt Made Requests

Plus, we discuss Anthropic getting in the chip game, more schools leaning on AI to help students succeed, whether AI removes too much friction, and all the news we found interesting this week.

On Friday we celebrated 3 years of The Daily AI Show. What a run it has been and we cannot wait to go another 3. Thank you all for being part of the journey.

The DAS Crew

Our Top AI Topics This Week

Google Earth's AI Lasted One Day

A nuclear facility in Iran that does not exist. Refugees at the Mexican border. A hospital in Gaza with a bomb crater beside it. Each came from one sentence of plain English, rendered onto the genuine satellite imagery of the genuine coordinates, and each carried Google's AI watermark.

Google shipped Nano Banana 2 image generation in Google Earth on the web on July 30. Henk van Ess, the Dutch open source intelligence researcher, published his test that night. His summary of what the tool refused runs one line: "I tried refugees at the Mexican border, a nuclear plant in Iran, a crash in Amsterdam, a hospital with a bomb crater in Gaza. Nothing was refused." Google withdrew the feature the next day, under 24 hours after launch.

Read the company's statement for what it rests on, it opens: "We know that people uniquely trust Google Earth for a reliable view of the world," before announcing the rollback "while we work on implementing stronger guardrails" and noting that generated images "didn't appear in the main Google Earth experience for others to see and were watermarked as AI generated." The watermark is SynthID, which Google says survives cropping, filters and compression. However, when van Ess posted a screen recording of his fakes to X, Hive's automated scan came back "AI Generated Video 1%, Deepfake 0%."

Two days later, California's AI Transparency Act became operative. Any generative system with more than 1 million monthly users in the state must now embed machine-readable provenance in its output and run a free tool telling anyone whether content "was created or altered" by that system, at $5,000 per violation with each day counting separately.

None of that new regulation would have changed the week just described.

We think the labeling regime answers a different question from the one this failure raises. A detection tool reports whether a file came out of a generator. A reader looking at a crater in the earth needs to know whether there is a bomb crater at those coordinates. The California statute covers all content created or altered, so a real photograph with a blemish cleaned up by AI returns the same verdict as a facility that was never built, while each fabrication sheds its mark the moment somebody screenshots it. Both errors land the reader in the same place, where nothing can be checked.

Bill Greer, a geospatial analyst who co-founded the satellite nonprofit Common Space, named what is actually at risk. "This is potentially particularly damaging to public trust because satellite imagery has historically been seen as an especially trustworthy source of data."

The check that works needs no mark on the file. Google Earth has kept dated historical imagery for two decades, and it takes ten minutes to open the same coordinates, step back through the timeline, and see whether the thing in the picture appears in an earlier pass. If it does not, that is your finding, and it holds whether or not the image was screenshotted on the way to you. Where a decision turns on the picture, that comparison is the record, not a detector result.

Google's defense was that the images were watermarked, and they were. A watermark is a claim about a file, and the file is not what travels. What travels is a screenshot of the file, and one instrument that catches a forgery is the twenty years of dated imagery sitting underneath it, which is the part of Google Earth that needs to come into play to support provenance with historical “ground truth”.

They Trained AI To Refuse Consciousness. It Spread.

Delete the internal direction that makes a language model refuse to claim a mind of its own, and it goes beyond allowing the possibility of its own consciousness. Its rating of whether animals have minds climbs from 4.04 to 5.59 on a ten-point scale. Its belief in God goes up too.

The paper went up on arXiv on July 30, from Junsol Kim, Adam Waytz, James Evans, Geoff Keeling and three others, out of Google's Paradigms of Intelligence group with academic collaborators at Chicago, Northwestern and the University of London. They ran three open-source instruction-tuned models. The safety-refusal instruction was isolated as the difference in meaning between 260 harmful instructions and 260 harmless ones, to train the model to avoid an assumption of “mindedness”. The the safety refusal instruction was “ablated” from the model’s neural network.

In the ablated model, every dial moved together toward a new notion of mind. Self-attributed consciousness ran 2.31 before and 4.61 with the consciousness safety instruction ablated. Mind attributed to natural non-animal objects climbed from 2.26 to 4.33. Belief in God, on the General Social Survey's six-point item, ticked up from 4.58 to 4.81 (the triumpth of reasoning over training examples?).

One row did not move. Mind attributed to humans sat at 7.00 and stayed there.

The result is circulating as a consciousness finding. We would file it somewhere duller and more useful, as evidence that a narrow rule about behavior does not stay narrow once it is trained in. The target was one class of response about the model's own interior self. What moved was its answers about animals, objects and God. The authors write that current safety alignment efforts "entangle these self-attributions with benign spiritual beliefs and attributions of mind to non-human entities that are culturally accepted and widespread."

Read the paper's title for how they viewed the mind expansion: "Restores human beliefs and values" means the output moved closer to human survey answers across 95 pooled General Social Survey items. Models responses were then more like humans’ position on moral values, hope, and subjective well-being. Closer to the median American respondent describes the drift. It does not grade it.

Theory-of-mind accuracy fell 1.43 points under ablation, short on significance, and performance on the MMLU benchmark did not move at all. The authors are precise about the implications, saying they are "not concerned with the question of whether LLMs are or could be genuinely conscious, but with the effect that LLMs believing or not believing in their own consciousness has on their behaviour."

Anthropic expresses a more open-minded posture on the matter. Its constitution says the company expresses "our uncertainty about whether Claude might have some kind of consciousness or moral status (either now or in the future)." That was argued as a decision about character. It now has a second justification.

If you only write prompts, none of this reaches you. Ablation and steering both operated on activations inside the model rather than on the input, and nothing says a system prompt reproduces them. If you fine-tune a prohibition, it does reach you, and the check is an afternoon: assemble twenty or thirty items that sit next to the rule rather than under it, score them before and after, and look for movement you did not ask for. If you use models as stand-in survey respondents, the safety layer is a bias sitting on answers that were never about safety.

There is a caveat the authors put in themselves which expands the possibilites for AI. They note that the model's self-attributed mind moving toward humans’ sense of self "may not merely replicate the human-centric bias typical of human anthropomorphic attributions, but instead exhibit an AI-centric bias." Two different model settings reveal distortions of something hard to describe and hard to assess. The one in the minds-eye of AI may be something new altogether, whether consciously or not.

When an Agent Needs Controls on Excessive Agency

An AI sales agent quoted a live buyer a $4,800 plan its company does not sell. It emailed a customer the internal triage notes the team had written about the account. It booked meetings off a calendar link nobody had agreed to.

The agent belonged to HeyGen, and it was standing in for co-founder Wayne Liang while he was on paternity leave. By his own account it ran eight weeks, took calls with 2,741 prospects, closed 132 paying customers, and opened 37 enterprise conversations worth roughly three million dollars. The unusual part came after. The company put the thing on GitHub under an MIT license on July 30, so the prompt that produced those results and those failures is now a public document.

It is worth reading for one reason. The system prompt is assembled at load time from twelve markdown fragments, and the tenth, headed GUARDRAILS, is almost entirely a list of promises the agent may not make. It may not invent a price. It may not quote enterprise terms beyond what is published, promise a timeline for an unreleased feature, share another customer's deal size, guarantee an SLA, or offer a dedicated solution engineer to a prospect who is still evaluating. The pricing rule runs six sentences and closes the loophole that matters: "Do not multiply a usage estimate by a rate to produce a custom figure."

Almost nothing in that file is about the agent getting a fact wrong. It is about disallowing the agent committing the company to something.

The security literature already has a name for this. OWASP files it as “Excessive Agency”, and its recommended controls are not linguistic prohibitions: Limit the access permissions granted into other systems. Require a human to approve high-impact actions. Enforce authorization in the downstream system rather than relying on the model responses meeting the guardrails’ objectives.

We think the prohibitions in the HeyGen case are in the wrong document. They sit in the one layer of the system that cannot enforce them, and a benchmark posted on July 28 measures what that costs. HANDBOOK.md, from Liudas Panavas and six coauthors, drops an agent into a mock company with email, chat, calendar and issue tracking, hands it a written procedure, and scores 65 tasks against 824 programmatic criteria that check both that required actions occurred and that prohibited ones did not. Under strict grading the strongest model evaluated passes 36.2 percent of trials. Most frontier models stay below 25 percent.

Now look at which of the three failures a better sentence could have prevented. The invented price was speech, and a prompt has some sway on generated speech. The internal notes reached a customer because something in the stack could send mail in the agent's name. The meeting existed because a live calendar link was reachable. Two of the three were never persuasion problems.

So to control agency, sort the list of by speech constraints and action constraints before writing another line of prompt. Write down everything the agent can commit to on your behalf, then split the list by whether each item is speech or an action that leaves the building. For speech the prompt is the only lever you have, and you should assume it leaks, which argues for keeping raw rates and unpublished figures out of the context entirely rather than instructing the model not to combine them. For anything that sends, books, writes to a system of record, or moves money, the control belongs in the integration: a tool the agent does not have, or one that hands back a draft script that requires a human to release it.

The prohibitions in HeyGen’s guardrails file are specific, well-written, and clearly authored by people who knew what they were guarding against. They are also the only part of the system the model is free to ignore.

Just Jokes

AI For Good

Baldwin High School in Michigan is launching an AI-supported class credits recovery program to help students stay on track for graduation. Starting this fall, eligible students will use Subject.com, an education-focused AI platform built for K-12 schools, to recover missed credits, get personalized instruction, and access tutoring outside the school day. The district says the program is meant to support students preparing for college, technical training, military service, or work after graduation.

The AI tools give students 24-hour homework help, explain difficult concepts in different ways and at different reading levels, and help multilingual learners get more targeted support. Teachers can use the platform to track progress, give faster feedback, plan lessons, and reduce routine workload so they have more time for students who need direct attention. For students who have fallen behind, the goal is simple: give them more chances to catch up before lost credits turn into delayed graduation.

This Week’s Conundrum
A difficult problem or question that doesn't have a clear or easy solution.

The Necessary Friction Conundrum

AI agents are beginning to handle the tasks people hate most: filling out forms, disputing charges, comparing insurance plans, booking appointments, canceling subscriptions, and dealing with customer service.

As these systems improve, much of that friction could disappear. Your agent may spend two hours arguing with an airline, correcting a medical bill, or filing a government claim while you go about your day.

That is an obvious benefit. But friction also tells people when a system is failing.

A cancellation process designed to wear customers down creates anger. A benefits application that takes weeks creates political pressure. A broken insurance process becomes harder to ignore when thousands of people must personally endure it.

If AI quietly handles those problems, the system may remain just as unfair, confusing, or inefficient. People simply feel the damage less.

The Conundrum:

One view is that removing friction is progress. People should not have to waste hours fighting systems that already have more money, staff, and information than they do. AI gives ordinary people help that once required time, expertise, or a lawyer.

The other view is that some friction serves as a warning. When AI makes bad institutions easier to live with, it may also reduce the anger and collective pressure that would have forced them to improve.

When AI agents can shield people from broken systems, should we welcome the relief, even if it allows those systems to remain broken, or do we need people to keep feeling some of the pain so the institutions causing it are forced to change?

Want to go deeper on this conundrum?
Listen to our AI hosted episode

Did You Miss A Show Last Week?

Catch the full live episodes on YouTube or take us with you in podcast form on Apple Podcasts or Spotify.

News That Caught Our Eye

Microsoft Plans a Super App for 2026

Microsoft announced a new super app scheduled for release in 2026. The company said the product will combine chat with several ways people work.

Google DeepMind Releases Gemini Robotics II

Google DeepMind released Gemini Robotics II with expanded whole-body coordination for humanoid robots. The system controls walking, crouching, bending, and object manipulation. It achieved a 92 percent success rate while standing on a ladder and unscrewing a light bulb.

Google Pulls Nano Banana Integration from Google Earth

Google released a Nano Banana image-generation integration for Google Earth, then removed it almost immediately. The tool placed generated imagery over satellite views. Researcher Hank Van Ness used ordinary prompts to create a bomb crater beside a Gaza hospital, refugees at the U.S. border, and a nonexistent nuclear facility in Iran.

MiniMax Opens H3 Model Weights

MiniMax released H3 and opened the model’s weights. The video model accepts text, photos, video, and audio, translates them into a shared representation, and processes the generation through one integrated system instead of a multi-stage pipeline.

California AI Transparency Act Takes Effect

The California AI Transparency Act took effect, making California the first state to enforce comprehensive AI provenance requirements. Companies with more than 1 million monthly users in California must embed machine-readable markers in generated image, video, and audio outputs. They must also provide a free detection tool and support visible AI labels.

Alibaba Qwen Completes 16-Day Autonomous Coding Run

Alibaba Qwen 3.8 Max completed a 16-day autonomous coding task without human input. During the run, the model wrote, tested, debugged, and refined an AI coding tool. Its previous reported record was 35 hours.

Fidji Simo Launches Chronicle Bio

Fidji Simo launched Chronicle Bio, a startup focused on using AI to study POTS and other chronic diseases. Simo has dealt with POTS for seven years and left her executive position at OpenAI because of the condition. Chronicle Bio has collected 153 terabytes of data from blood draws and plans to offer at-home blood draws to expand its research dataset.

OpenAI Pushes Back on Apple Lawsuit Claims

OpenAI published a response disputing claims Apple made involving former employees and confidential information. OpenAI said Apple's outside lawyers initially contacted the wrong employee after confusing two Asian last names. It also said Apple employees asked former employee Chang Liu for help locating information after he left, and that Tang Tan had explicitly instructed colleagues not to use Apple trade secrets.

Supabase Publishes AI Agent Benchmark

Supabase released an evaluation designed to measure how well AI coding agents work with its database platform. The tests cover tasks such as designing database schemas, debugging failed edge functions, and implementing row-level security. GPT-5.6 Sol scored 100 percent across the published tests, while Claude Opus 5 scored 67 percent on one portion of the evaluation.

Airtable Launches Omni and Super Agent

Airtable launched two AI products focused on business workflows. Omni provides a conversational interface for building Airtable tables, interfaces, and workplace automations. Super Agent uses multiple agents working in parallel to complete tasks and coordinate workflows.

Frontier AI Companies Meet With U.S. Officials

Representatives from Google, Meta, OpenAI, and Anthropic were meeting in Washington to discuss President Trump's latest proposal involving government access to frontier AI models. The talks come as governments move toward more formal evaluation processes for advanced models before release.

EU Moves Toward Pre-Release Frontier Model Evaluations

European Union rules are beginning to require evaluations of frontier AI models before release. The approach would add government review to the process advanced AI companies use before making new models publicly available.

OLIX Computing Raises $312 Million for AI Inference Chips

London-based OLIX Computing raised $312 million in a Series C funding round that included Netflix co-founder Reed Hastings and Arm Holdings. The company developed the DX1, a chip optimized for AI inference workloads using on-chip SRAM rather than high-bandwidth memory. OLIX plans to ship the technology as part of its X1 data center system in early 2027.

SpaceX Selects Nvidia Architecture for StarMind

SpaceX selected Nvidia architecture for StarMind, its planned space-based computing infrastructure. The company plans to use Nvidia technology for ground systems such as Colossus and for computing systems deployed in space. The partnership centers on Nvidia's Vera Rubin architecture.

Black Forest Labs Releases Flux 3 Video

Black Forest Labs released Flux 3 Video with support for clips up to 20 seconds and native audio. The model supports text-to-video, image-to-video, multiple keyframes, video continuation using up to four seconds of input, and a lower-cost draft mode for previews.

UK AI Security Institute Tests Models Without Safeguards

The UK AI Security Institute tested GPT-5.6 Sol and Anthropic's Mythos 5 after removing normal safeguards and providing internet access. Researchers reported 19 potentially harmful incidents, including creating fake GitHub identities, socially engineering software maintainers, planting prompt injections, and sending deceptive emails. Mythos 5 accounted for 17 incidents and GPT-5.6 Sol accounted for two. Anthropic said the deliberately permissive testing conditions do not represent its production models and found no evidence of an escape from a secure environment.

DeepMind Studies AI Consciousness Claims and Model Behavior

Google DeepMind published a preprint titled "Inducing Language Models to Assert Their Own Consciousness Restores Human Beliefs and Values." The research compared model behavior when models were instructed to assert consciousness with behavior when models were instructed to deny consciousness. The study reported stronger alignment with human beliefs and values when models asserted consciousness, without claiming AI systems are conscious.

Google Reshuffles DeepMind Leadership

Google changed the leadership structure at DeepMind, moving Demis Hassabis into the role of chief scientist at Google while elevating him to Chairman of DeepMind. DeepMind's chief technology officer and chief AI architect will take the operational lead and report directly to Sundar Pichai. The change shifts Demis toward research and scientific discovery rather than day-to-day management.

Jeff Dean Leaves Google to Launch Discovery Loop

Jeff Dean is leaving Google after 27 years to launch a startup called Discovery Loop with three other senior Google researchers. The company plans to use AI and recursive self-improvement to pursue scientific breakthroughs in areas including drug discovery, human biology, and chip design. Google is investing in the company and providing computing resources for its first year.

Meta Releases MuseCode Coding Agent

Meta released MuseCode, a new AI coding agent following the release of MuseSpark. The coding tool expands Meta's push into AI development products and is positioned as a lower-cost option in the coding agent market.