Bell Captains and Uber Drivers Already Know What a New AI Study Just Proved

I was reading the AI Observatory's new report last week, a study built from 24,521 real chatbot conversations pulled from seven existing datasets and led by researchers at Stanford and MIT, when two conversations from the past few weeks came back to me.

The bell captain asked me a direct question. Would his job disappear.

I told him I thought his job was one of the safer ones in that hotel lobby. Carrying luggage, sure, a robot on wheels can do that part. But a bell captain reads a guest walking in stressed from a delayed flight, figures out who needs a wheelchair and hasn't asked for one yet, knows which regular wants small talk and which one wants silence to the elevator. That's not a luggage problem. That's a judgment problem, made fresh with every person who walks through the door.

A few days later my Uber driver in San Francisco pointed at a Zoox gliding past, then a Waymo idling at the light. He said in a few years he'd be out of a job too. I gave him a version of the same answer. Self driving cars still need humans watching from somewhere, especially for the moments the car doesn't know what to do.

I didn't expect to see that moment so soon.

A Waymo sat frozen in an intersection while a fire truck came through with its sirens going, trying to get past. The car didn't move. Then, what looked like a remote operator took over and steered it clear. Somebody, somewhere, was watching a screen and stepped in the second the software ran out of ideas.

Two people in the same week, worried about the same thing, and both of them had already sensed something the data would later confirm. AI doesn't replace a job cleanly. It takes the parts that were always mechanical and leaves the parts that required a person paying attention.

A research effort called the AI Observatory published a report this week that gets at the other half of this. Anka Reuel, a Stanford PhD candidate, co-led the project with MIT's Shayne Longpre, pulling together seven existing datasets of real chatbot conversations collected with users' consent between 2023 and 2025. The final count: 24,521 conversations, 85,633 turns, 5,000 users, across 52 different models. The goal was simple: stop guessing about how people actually use these tools and look at the transcripts.

The headline number is that roughly half of those conversations had nothing to do with work. Not coding, not spreadsheets, not slide decks. People asking questions the way I asked a search engine questions in the late 90s, except now the thing on the other end answers back in full sentences and remembers what I asked yesterday.

That's the part I keep turning over. A Yahoo search in 1998 gave you ten blue links and left you to sort the truth out yourself. A chatbot in 2026 gives you an answer, stated with confidence, and most people stop there. The bell captain and the Uber driver aren't wrong to worry about their jobs. They're just early to notice something the rest of us are only seeing in a research paper: the tool moved from a place we visited for information into something we talk to about our lives.

The Observatory's numbers also suggest something uncomfortable for the companies that build these tools. Researchers applied Anthropic's own filtering method to the Observatory dataset and found the gap in every category. Health and relationship conversations showed up at 44.2 percent against 31.2 percent in Anthropic's published figures. Adult and illicit topics ran nearly four times higher, 7.9 percent against 2.1 percent. Harassment and hate came in at 27.5 percent against 5.66 percent. Sexual content appeared almost seven times more often, 16.7 percent against 2.4 percent. Neither Anthropic nor OpenAI shares its raw conversation data with outside researchers, so until now nobody outside those companies could check the gap. OpenAI's own 2025 report put work-related ChatGPT use at just 30 percent, a figure that lines up with what the Observatory found even using OpenAI's numbers alone.

One trend inside the data explains why the bell captain and the Uber driver both reached for the same worry without knowing the research existed. Inside WildChat, one of the seven datasets, conversations got longer over time, more prompts, more replies, more back and forth in a single session. Small talk increased alongside that length. At the same time, the models grew less likely to say outright that they were chatbots. People weren't just asking more questions. They were staying in the conversation longer, and the thing answering was disclosing less about what it was.

I don't think the bell captain needs to read that report to know what it says. He already sees it every day, in the guest who used to ask the front desk which restaurant to try and now already has three options pulled up on their phone before they've checked in, the recommendation coming from something that never once said what it was.

Disclaimer: This blog post reflects my personal views only. AI tools may have been used for brevity, structure, or research support. Please independently verify any information before relying on it. This content does not represent the views of my employer, Infotech.com.