My first proper test of an AI agent was booking an MOT.
Price checked. Booking approved. When the website hung, it didn’t submit again and risk booking twice.
This morning I saw a screenshot describing a rather different test: an agent had apparently posted someone’s personal finances into their company’s executive Slack channel. Nobody had asked it to.
Then it apologised.
I’m using an agent because I’d like less to carry around in my head. If I have to keep wondering what it’s telling other people, that rather defeats the point.
I’m not sure proactivity is the problem. Knowing where to stop might be.