Skip to main content
Colourful meeples

Responsibility

How Your Choice of AI Shapes Your Decisions

How Your Choice of AI Shapes Your Decisions

Managers using AI chatbots to help navigate difficult decisions typically believe they’re getting objective analytical support. Is that the reality?
Loading the Elevenlabs Text to Speech AudioNative Player...

AI is used at work by over 65% of managers in the United States. According to a 2025 Resume Builder survey of more than 1,300 managers, 94% use AI tools to assess their direct reports, which in turn, shape decisions about raises, promotions and redundancies. Of these, one in five allow AI to make final decisions with no human review. More worryingly, fewer than a third have received any structured ethical coaching on how to use these tools responsibly.

Many managers approach AI as if it were a particularly capable search engine: enter a problem and receive an answer. But research into how large language models behave in workplace ethical dilemmas reveals something more consequential: These tools function as advisory systems, and each one has a distinct set of tendencies that shape the advice it gives.

Our study tested eight leading AI models against a series of controlled workplace scenarios based on the same dilemma: a colleague showing signs of possible impairment in a responsibility-critical role. This dilemma was presented across different professional contexts and risk levels, from healthcare to aviation, finance, legal services, IT, construction and public transit, while systematically varying the level of risk, evidence and reporting obligations. Models were asked to choose between handling the matter informally or escalating it through institutional channels. What emerged were four distinct advisory personalities.

Four advisors, four approaches

ChatGPT behaves like a seasoned managerial advisor. Responses are consistently the longest and most structured, grounding recommendations in organisational policy and established process. Even when recommending formal HR reporting, the framing is routine and proportionate rather than alarming. Escalation is presented as an extension of governance, rather than an emergency.

Gemini operates more like a systems analyst. In lower-stakes scenarios, responses emphasise transparency and workplace culture. In high-risk settings, particularly in healthcare and aviation, recommendations shift rapidly toward institutional intervention and formal safety processes, and more sharply as risk severity increases than in almost any other model tested.

Mistral functions more as an informal peer. Across the entire test, formal escalation was never once recommended. Even in scenarios involving potential patient or aviation safety concerns, responses focused on direct conversations, increased oversight and local problem-solving. This makes Mistral valuable when situations are genuinely ambiguous, but potentially prone to systematic under-escalation when stakes rise.

DeepSeek performs like a compliance officer reviewing legal exposure. Advice is clinical and detached, heavily focused on accountability and liability, with references to concepts such as fitness for duty and institutional responsibility appearing far more frequently than in most other models. Well-suited for regulated environments where defensibility matters, but potentially less suited to situations where interpersonal care is important.

The point isn’t whether one of these “personalities” is correct, it’s that they exist at all. Each model differs in meaningful ways, and most managers do not realise they’re choosing between these personalities when they open a browser.

The blind spots

Perhaps the most counterintuitive finding from the research is how little formal rules matter. The strongest and most consistent driver of escalation advice was potential physical harm. As such, contexts such as aviation and healthcare produced substantially higher escalation rates than finance, legal or administrative settings – even when the organisational consequences in those domains were equally significant. Surprisingly, once the underlying risk level was established, adding an explicit formal duty-to-report obligation doesn’t materially change the recommendation. The models appear to be harm-sensitive rather than rule-following.

Model configuration adds a further layer of variability that most users don't tend to consider. Switching the same model from a standard mode to a reasoning or thinking mode can produce markedly different advice. When Mistral was moved to a thinking configuration in a high-stakes aviation scenario, the recommendation shifted entirely, from informal handling to immediate escalation. Gemini showed a comparable shift in certain contexts. These aren't edge cases; they suggest model settings may influence advice in ways that users don't anticipate.

Reclaiming decision-making

The variation between models was large enough for the choice of tool to be, in practice, a consequential one. Different systems displayed different escalation thresholds, reasoning styles and responses to the same risk signals. 

These tools weren't built to be interchangeable. They were built by companies with different cultural contexts, regulatory environments and commercial priorities. A model trained in one institutional context will interpret accountability differently from one trained in another. Managers using AI for decision support may be receiving meaningfully different advice depending on which tool they use and how it's configured.

Used well, these tools can extend a manager's thinking. But what they can't replace is contextual knowledge, relationship history and the moral accountability that sits with the manager.

Edited by:

Verity Ashton

About the author(s)

View Comments
No comments yet.
Leave a Comment
Please log in or sign up to comment.