Left unsaid: The bot wasn’t very confident in its own advice
Launching into young adulthood involves grappling with a steady stream of big decisions that can be life altering. Deciding whether to opt for college or vocational school, or perhaps enlist in the military, hits right around age 18. The ensuing decade or so tends to bring its own cascade of firsts — whether to get married, start a family, buy a home.
That none of these decisions has an objectively correct answer adds to the challenge, as young adults have to figure out what is best for them. Friends likely facing the same questions can provide a willing ear but are rarely more expert, and live in the same echo chamber.
Artificial intelligence is a new sounding board for young adults weighing life-altering decisions. Surveys already show that Gen Z is turning to large language models not just for work tasks but for more personal guidance. More than 40% of Gen Z ChatGPT users have consulted the LLM for career advice.
A working paper reports that when ChatGPT took on the role of advice-giver to young adults mulling consequential decisions — go to college or vocational school, join the Army, buy a house, etc. — its recommendations often differed from what more than 300 people advised when asked the same questions. UCLA Anderson’s Kathleen Ngangoue and New York University’s Andrew Schotter and Bill Wang suggest that’s potentially helpful. A bot whose recommendations don’t merely mimic human advice widens the range of possibilities a young adult might weigh.
Notably, the researchers observed through repeated testing that the chatbot was extremely consistent in its advice. Asked 100 times to advise the same person on a given decision, the bot typically spit out the same recommendation. Anyone consulting the bot in this way might reasonably presume repetition is a sign of confidence. Yet the researchers tease out an important nuance: Consistency is not to be confused with conviction.
An important structural feature of the research is that the bot was limited to a binary choice for each topic: Recommend X or Y. Within that structure, the researchers had the LLM rate its confidence in its own advice on a scale of 0 to 10. The median was 6. That gap between steady answers and middling confidence reflects an important feature of these systems. The model doesn’t settle on an answer so much as rank the possible ones, scoring each on how likely it is to be the best answer, the highest scoring possibility is what emerges as the advice. Yet when two answers score nearly the same, one can still come out on top again and again.
And ChatGPT pretty much acknowledges the disconnect. The researchers fed the written explanations ChatGPT had given for its advice back to the LLM and asked it to self-evaluate its justifications. The bot read those explanations as more emphatic and more confident than its own numerical scores had been.
“Users who are unfamiliar with the internal workings of LLMs should be cautious not to interpret highly resolved recommendations across repeated prompts as highly confident ones,” the authors write.
Income Impacts Advice, Not Race or Gender
To explore how the bot processes information and lands on advice, the researchers created fictional advisees that varied by race, gender and income — a wealthy Black man, a low-income white woman, etc. Because ChatGPT-5 is designed to avoid acting on race and gender directly, those traits were signaled through names such as Jamal and Laquisha for Black advisees, Jack and Emily for white ones. Testing confirmed ChatGPT read the names as intended.
The researchers also had the bot role-play advisors with similar demographic backgrounds. The characteristics of advisors made little difference in the advice that was doled out. What mattered was who was receiving the advice, not who was giving it.
On half of the 10 decisions the bot gave everyone the same advice. On the other half its advice varied — and income, not race or gender, was what overwhelmingly mattered.
The starkest divide showed up in three decisions. Low-income advisees were steered toward vocational school, told to stay employed rather than start a business, and strongly encouraged to enlist in the military. Higher-income advisees were pointed toward college, encouraged to consider self-employment and advised against the Army.
Divorce was a notable exception where the difference in advice wasn’t tied to income: Women and white advisees were somewhat more likely to be advised to leave a marriage. Insurance advice also varied, but without a similarly clear demographic pattern.
The humans surveyed through YouGov offered different advice. On the decisions where ChatGPT gave rich and poor advisees different recommendations, people tended to give much more similar advice. And on decisions where the bot had given rich and poor similar advice, real people tended to serve up different advice based on income.
In other words, the bot disagreed with people about which decisions even called for different advice. Ngangoue and her colleagues suggest this is evidence that AI advice is its own channel of guidance rather than a digital echo of human advice.
Bot Versus Real World
The researchers then set the bot’s recommendations against what people in each group actually do, drawing on Census data. They found a pattern in which, even when the bot gave everyone the same advice (don’t buy a house at 24 regardless of income), it lands differently based on where someone sits in the real world. Telling a poor 24-year-old to not buy the house confirms what the Census data shows poor 24-year-olds already do. But to tell a wealthy 24-year-old not to buy a house runs counter to what wealthy 24-year-olds actually do. The pattern held across other scenarios: The bot reinforced choices already common among lower-income people while nudging higher-income people away from their group’s norm.
The reasons ChatGPT offered also differed by income. Averaged across the decisions, its justifications for high-income advisees put more emphasis on long-term benefits, potential regret and happiness. For low-income advisees, they put more weight on short-term benefits, risk and ease of implementation. While the mix varied by the decision at hand, aggregated across decisions, this income pattern remained.
In the figures below, the X marks how likely the LLM is to recommend a given choice; the dot shows how often that group actually makes it. Close together, the advice confirms what the group already does. Far apart, it runs counter. Income is the dividing variable in every scenario except divorce, where gender mattered most.
A Different Set of Choices
Income’s role showed up in starker form when the researchers dropped the either/or format and asked ChatGPT to name colleges a high school student might consider.
Across repeated runs, students with identical academic records but different family wealth were fed different lists of schools.
Recommendations for higher-income students clustered among institutions ranked in the top 50 globally, suggesting the advice was based on academic ability. The lists for lower-income students were more scattered, reaching well past the top 100, and importantly they also included schools that provide need-blind financial aid. That suggests that when free to generate its own advice (and not be shackled by deciding between two possible choices) the bot resorts to income profiling.
That’s not proof of bias, but rather an algorithmic outcome from a bot trained on existing data. Yet however implicit or unintentional, it suggests that AI-generated advice may (inadvertently) perpetuate existing norms that limit opportunity.
And as for the fact that the bot gives different advice than humans on certain consequential decisions, that may be both a feature and a bug. Being exposed to more possible paths to take can be helpful. Yet like all consequential advice that is highly personal, users would benefit from understanding that consistency in advice is not necessarily a signal that the bot has a high level of confidence in what it’s recommending.
Featured Faculty
-
Kathleen Ngangoué
Assistant Professor of Global Economics and Management
About the Research
Ngangoue, K., Schotter, A., & Wang, B. (2026). Life’s Big Decisions With AI Advice.