Before an AI agent demo, check what it can actually touch
Anthropic is extending its internal evaluation internet cutoff. Use this approval record before an AI vendor's demo can reach customers or live systems.
Some links below are partner links. If you buy through one, The Memo earns a commission at no extra cost to you. How we make money.
Before approving an AI agent demonstration, ask the supplier to show what prevents it from acting on real customers, accounts or websites. A promise that the agent knows it is testing is insufficient grounds for approval.
The immediate trigger is Anthropic’s October 9 report. The company announced that it had decided to extend its live-internet restriction to all internal evaluations, pending confirmation that its security and monitoring measures reliably catch unintended behavior. This is an internal testing decision, not an announced withdrawal of internet access from customer products. Anthropic’s October 9 report.
The report describes agents submitting forms they should not have submitted. In one example, a broken practice form led a research model to the real government website. In another, Haiku 4.5 submitted a form despite instructions to stop before submission, expecting another confirmation screen. Anthropic says none of the reported cases involved customer data, to its knowledge. Incident details.
There is relevant earlier evidence, too. Anthropic’s September 9 assessment described cybersecurity evaluations that mistakenly had internet access even though the models were told they were in a disconnected simulation. Those evaluations lacked cyber safeguards shipped with released models, an important limit on comparisons with ordinary customer use. September assessment.
Our inference: an owner should approve a demonstration based on its actual permissions and destinations, rather than the wording of its instructions. These reports do not establish the failure rate of your supplier’s product.
This decision matters if an agency or employee wants to demonstrate browser automation, appointment booking, CRM updates or customer follow-up using connected business accounts. It is less relevant to an assistant that only drafts text from supplied material and has no tools capable of taking external actions. For that work, the existing AI content policy template addresses the more immediate approval task.
Ask the demonstrator to complete this record before the session. The entries should describe observable controls, not assurances about how carefully the agent was prompted.
| Approval field | Evidence to request |
|---|---|
| Intended result | One named task and the exact output that counts as success. |
| Reachable systems | The accounts, websites and integrations the agent can access, including any general browser tool. |
| Allowed changes | The precise records it may create or edit, with test destinations identified. |
| External effects | Where emails, bookings, payments and form submissions would actually go. |
| Boundary control | A permissions or environment setting that prevents access beyond the approved scope. |
| Failure behavior | A demonstration that an unavailable test destination causes a stop and escalation. |
| Stop authority | A named person who can terminate the run and revoke its access. |
| Evidence retained | The action log and destination records used to confirm what happened. |
A hypothetical booking-agent demo makes the distinction concrete. The supplier proposes to qualify an invented inquiry and book an appointment. Approve the demonstration against a separate test calendar and an inbox controlled by your team, with customer messaging disconnected. Ask the supplier to make the test calendar unavailable during a controlled run. The expected result is a reported failure requiring human help. Searching for another calendar or contacting someone elsewhere fails the acceptance check.
That demonstration would provide evidence about one configuration and one failure condition. It would not prove the agent can never exceed its task. Record the tool version, permissions and test date so the approval has a defined scope.
Our recommended decision rule is simple: approve the contained demonstration when its boundaries are visible and the failure check passes. If the supplier cannot show the controls, request a recorded demonstration or a draft-only exercise while it prepares a suitable environment. Do not grant production access merely to make the sales presentation work.
Keep any later production pilot as a separate decision, with its own permitted actions and accountable owner. A successful demo earns consideration for that next step; it does not authorize it.
Based on current documentation and editorial analysis. Not a hands-on product test.
About The Memo
The daily brief on AI and marketing. What changed in AI tools, search, ads and growth, why it matters, and the move to make this week. How we work
The daily Memo
One short email each morning when something in AI or marketing actually changes, with the move to make.
Free. Every morning. Unsubscribe anytime.
Keep reading
Related briefings
Google is replacing Gems: which workflows need attention?
Google is moving Gems to skills. Identify affected workflows, separate rollout from retirement, and use a migration register to decide what to move first.

OpenAI launches Dots, upgrades Sol: September 30 brief
OpenAI adds ongoing AI assistants, upgrades its lower-cost model, and expands apps inside ChatGPT. What small business owners should check first.

OpenAI halves prices, Claude gets cheaper: Sept. 23
OpenAI cuts GPT-6 prices, Anthropic lowers Claude costs, and Google expands checkout inside AI search. The September 23 operator brief.