whynotroman

News · Oct 09, 2026 · 7 min

How to work with AI, October 9. Five tips: ask for proof and set a spending limit

Five tips from Anthropic and Simon Willison in simple words. How to ask AI for proof, check numbers in a dashboard, turn on reasoning for math, set a spending limit and limit access. For each tip I show how to use it in car sales.

Today I chose five tips that you can use the same day. Three of them are about not taking AI at its word. Two are about money and access, where a mistake costs a lot.

I added my opinion to each tip. How it works in car sales and in a talk with a person who chooses a car.

Tell the agent what proof shows the task is done

The Anthropic team published a guide to cloud sessions in the Claude Code blog. Its main tip is about how you write the task. Say what is wrong, what the result looks like and how the agent will prove it. One task per session.

In the example, the agent had to fix a flaky test. The proof was to run the test suite at least 30 times in a row. It ran the suite 40 times with no failures. In another session the agent started the server and sent a request to every address. This found five errors that you cannot see by reading the code.

Check the result against the proof first, and read the agent's summary after.

In simple words. Do not ask "do it". Ask "do it and show me how you checked".

My opinion. A customer asks about a model with a loan. If you only ask for an answer, the payment comes from nowhere.

I would write it this way. "Answer the customer about a loan for this model. Take the rate only from the attached document. At the end show the payment step by step and name the line of the document with the rate."

Then I count one option myself. AI prepares, and the person decides and answers for the number. The price in an email is your promise to the customer.

Source Claude Code blog

Click a number in the dashboard and read how it was calculated

On October 8, Anthropic released Claude Dashboards in beta. A dashboard is built from a plain-language question and updates with the data. It connects to warehouses like BigQuery, Snowflake and Databricks, and to the CRM Salesforce.

Every number opens. You click a figure and see the query that calculated it. A chart shows when it was last refreshed. Dashboards are on paid plans. For company admins the function is off by default.

In simple words. Do not trust a nice number. Open how it was made and check one row by hand.

My opinion. A sales manager asks for a dashboard of "time to first reply by lead source". The number looks good. I would click it and check whether night leads and phone leads are in the count. If they are not, the average reply looks faster than it is.

Then I would list the managers whose reply takes over an hour. AI does not replace the manager. It shows who did not do the work. Car DMS systems are not among the connections, and the first question for the vendor is about the export from the DMS.

Source Claude

Turn on reasoning when you ask AI to calculate

Simon Willison, a developer who writes a blog about AI, tested the Qwen 3.8 model on adding numbers. The model ran on his own computer, and the answer had to be written in words. It repeats a test that Colin Fraser did on GPT-4o two years ago.

Without reasoning the model was right on 23.6% of the examples. With numbers of one to three digits it was right 97% of the time, and with numbers of 10 to 13 digits only 6.4%. With medium reasoning turned on, 167 of 169 examples were right. This was a pilot with one example for each number length, so Willison expects different numbers on a repeat.

In simple words. Without "think", AI writes digits from memory and fails on long numbers. With reasoning it counts step by step.

My opinion. A loan payment, a residual value, the gap in a trade-in price. These are the calculations where a model without reasoning writes a nice mistake with confidence. In ChatGPT and Claude I would turn on the reasoning mode every time the request has money in it.

I would also add a phrase to the request. "Show the calculation step by step." If there are no steps, do not trust the number. One pilot run is no guarantee, so I always keep one check by hand.

Source Simon Willison's Weblog

Set a hard spending limit where an agent works

On October 3, Willison wrote that paid services need a hard spending limit by default. An email warning does not stop spending. A limit returns an error when the monthly budget is used up. In his view, agents make spending easy to reach, and the bill grows while the owner sleeps.

He points to a trend. Amazon Web Services launched spending limits in September, and a project pauses until the end of the month. The feature is not yet open to all customers. Google Cloud added Spend Caps in July. Willison wants agents to recommend services with a hard limit.

In simple words. It is better to get a refusal at the end of the month than a bill of thousands of dollars in the morning.

My opinion. A dealer connects an AI agent to night leads and pays for each reply. If a dialogue loops or someone starts a mass mailing, the bill grows with no call. I would ask the vendor one thing. "What monthly spending limit is set for us, and what happens when it is reached?"

A good answer sounds like this. The agent stops, I get an email, and customers get a message that a manager will reply in the morning. A bad answer sounds like "we have no limit".

Source Simon Willison's Weblog

Open only the access the agent needs to do the job

The same Anthropic guide describes levels of network access. "None" for untrusted code, "Trusted" for common repositories, "Custom" for private services and "Full" only when nothing else works. Environment variables are visible to everyone who uses that environment, so you should not keep secrets there.

There are two more facts in the guide. Anthropic stores the session transcript, and the period depends on the plan and the model-improvement setting. Even with access set to "None", requests to the Anthropic API still go out, so information can leave the environment.

In simple words. Give the agent exactly the rights the task needs. Close the rest, and do not put passwords where others can see them.

My opinion. In a dealership this is about the table with customer phone numbers. I would upload a file without phones and surnames. To check an answer, a request number, a model and the text of the question are enough.

Then I would open the storage policy of the chosen service and write down one line. Where the chats are kept and for how long. If you do not know the period, do not upload customer details. Show that line to the lawyer before you start, not after a leak.

Source Claude Code blog

What to watch

These are new releases and talks from the last two days. I did not review them yet, and the description is from the title.

Next note · 10 minHow to work with AI, October 8. Remove extra instructions and test AI on your own questions →Six tips from the writers of Every, Anthropic, OpenAI and Simon Willison in simple words. How to cut instructions, how to test AI on real customer questions and how to work with Claude in Google Docs. For each tip I show how to use it in car sales.

Roman Spitsyn. I build sales teams where AI agents do the work and people close the deals.

Read next

← All notes