News · Oct 09, 2026 · 7 min
How to work with AI, October 9. Five tips: ask for proof and set a spending limit
Five tips from Anthropic and Simon Willison in simple words. How to ask AI for proof, check numbers in a dashboard, turn on reasoning for math, set a spending limit and limit access. For each tip I show how to use it in car sales.
Today I chose five tips that you can use the same day. Three of them are about not taking AI at its word. Two are about money and access, where a mistake costs a lot.
I added my opinion to each tip. How it works in car sales and in a talk with a person who chooses a car.
Tell the agent what proof shows the task is done
The Anthropic team published a guide to cloud sessions in the Claude Code blog. Its main tip is about how you write the task. Say what is wrong, what the result looks like and how the agent will prove it. One task per session.
In the example, the agent had to fix a flaky test. The proof was to run the test suite at least 30 times in a row. It ran the suite 40 times with no failures. In another session the agent started the server and sent a request to every address. This found five errors that you cannot see by reading the code.
Check the result against the proof first, and read the agent's summary after.
In simple words. Do not ask "do it". Ask "do it and show me how you checked".
My opinion. A customer asks about a model with a loan. If you only ask for an answer, the payment comes from nowhere.
I would write it this way. "Answer the customer about a loan for this model. Take the rate only from the attached document. At the end show the payment step by step and name the line of the document with the rate."
Then I count one option myself. AI prepares, and the person decides and answers for the number. The price in an email is your promise to the customer.
Source Claude Code blog
Click a number in the dashboard and read how it was calculated
On October 8, Anthropic released Claude Dashboards in beta. A dashboard is built from a plain-language question and updates with the data. It connects to warehouses like BigQuery, Snowflake and Databricks, and to the CRM Salesforce.
Every number opens. You click a figure and see the query that calculated it. A chart shows when it was last refreshed. Dashboards are on paid plans. For company admins the function is off by default.
In simple words. Do not trust a nice number. Open how it was made and check one row by hand.
My opinion. A sales manager asks for a dashboard of "time to first reply by lead source". The number looks good. I would click it and check whether night leads and phone leads are in the count. If they are not, the average reply looks faster than it is.
Then I would list the managers whose reply takes over an hour. AI does not replace the manager. It shows who did not do the work. Car DMS systems are not among the connections, and the first question for the vendor is about the export from the DMS.
Source Claude
Turn on reasoning when you ask AI to calculate
Simon Willison, a developer who writes a blog about AI, tested the Qwen 3.8 model on adding numbers. The model ran on his own computer, and the answer had to be written in words. It repeats a test that Colin Fraser did on GPT-4o two years ago.
Without reasoning the model was right on 23.6% of the examples. With numbers of one to three digits it was right 97% of the time, and with numbers of 10 to 13 digits only 6.4%. With medium reasoning turned on, 167 of 169 examples were right. This was a pilot with one example for each number length, so Willison expects different numbers on a repeat.
In simple words. Without "think", AI writes digits from memory and fails on long numbers. With reasoning it counts step by step.
My opinion. A loan payment, a residual value, the gap in a trade-in price. These are the calculations where a model without reasoning writes a nice mistake with confidence. In ChatGPT and Claude I would turn on the reasoning mode every time the request has money in it.
I would also add a phrase to the request. "Show the calculation step by step." If there are no steps, do not trust the number. One pilot run is no guarantee, so I always keep one check by hand.
Source Simon Willison's Weblog
Set a hard spending limit where an agent works
On October 3, Willison wrote that paid services need a hard spending limit by default. An email warning does not stop spending. A limit returns an error when the monthly budget is used up. In his view, agents make spending easy to reach, and the bill grows while the owner sleeps.
He points to a trend. Amazon Web Services launched spending limits in September, and a project pauses until the end of the month. The feature is not yet open to all customers. Google Cloud added Spend Caps in July. Willison wants agents to recommend services with a hard limit.
In simple words. It is better to get a refusal at the end of the month than a bill of thousands of dollars in the morning.
My opinion. A dealer connects an AI agent to night leads and pays for each reply. If a dialogue loops or someone starts a mass mailing, the bill grows with no call. I would ask the vendor one thing. "What monthly spending limit is set for us, and what happens when it is reached?"
A good answer sounds like this. The agent stops, I get an email, and customers get a message that a manager will reply in the morning. A bad answer sounds like "we have no limit".
Source Simon Willison's Weblog
Open only the access the agent needs to do the job
The same Anthropic guide describes levels of network access. "None" for untrusted code, "Trusted" for common repositories, "Custom" for private services and "Full" only when nothing else works. Environment variables are visible to everyone who uses that environment, so you should not keep secrets there.
There are two more facts in the guide. Anthropic stores the session transcript, and the period depends on the plan and the model-improvement setting. Even with access set to "None", requests to the Anthropic API still go out, so information can leave the environment.
In simple words. Give the agent exactly the rights the task needs. Close the rest, and do not put passwords where others can see them.
My opinion. In a dealership this is about the table with customer phone numbers. I would upload a file without phones and surnames. To check an answer, a request number, a model and the text of the question are enough.
Then I would open the storage policy of the chosen service and write down one line. Where the chats are kept and for how long. If you do not know the period, do not upload customer details. Show that line to the lawyer before you start, not after a leak.
Source Claude Code blog
What to watch
These are new releases and talks from the last two days. I did not review them yet, and the description is from the title.
- OpenAI. Introducing GPT-6 in ChatGPT with Intelligent UI, a demo of the new answers, 493 thousand views
- OpenAI. Build plugins for ChatGPT, 77 thousand views
- OpenAI. How OpenAI puts ChatGPT to work, a talk from DevDay 2026
- OpenAI. How Oracle uses ChatGPT in recruitment, a look at a work process
- AI Engineer. How to run AI agents safely in production, a talk by Viren Baraiya
- AI Engineer. Why AI agents should have their own sandbox, a talk by Philipp Schmid of Google DeepMind
Read next
- Oct 07, 2026
NewsHow to work with AI, October 7. Six methods for people who sell carsSix tips from GitHub, Ethan Mollick, Anthropic and Simon Willison in simple words. How to check an AI answer with a second model, how to argue with a model, how to run long tasks and how to give access to an agent. For each tip I show how to use it in a conversation with a car buyer.8 min - Oct 09, 2026
NewsAI and car sales, October 9. Cox builds Fullpath's AI platform into VinSolutions CRMCox builds Fullpath's AI platform into VinSolutions CRM, the EU prepares quotas on Chinese hybrids, Mercedes sells record electric cars, and Volkswagen sets aside 725 million pounds for car finance claims. Short news, my opinion, a link to the source.19 min