Skip to main content
Pink scales

Responsibility

The Other AI Is Academic Integrity

The Other AI Is Academic Integrity

Fabricated data, statistical sleight of hand, and now AI. Social science's system for catching its own mistakes is under more pressure than ever.
Loading the Elevenlabs Text to Speech AudioNative Player...

In 2005, Barbara Fredrickson, a respected psychology professor at the University of Michigan, published research “proving” that 2.9103:1 was the minimum ratio of positive to negative emotions required for people to flourish. Nine years later, 11 mathematical errors were found in the calculations behind it. Worse, it emerged that her co-author Marcial Losada had fabricated the underlying data. Worse still, the case turned out to be far from unique.

Why does research integrity break down and what can be done about it? These questions sat at the centre of the Workshop on Research Integrity, held on 29 and 30 April at INSEAD’s Europe Campus in Fontainebleau, chaired by Professor N. Craig Smith of INSEAD and Associate Professor Nina Strohminger of The Wharton School. The event was organised by the INSEAD Ethics and Social Responsibility Initiative (ESRI)with support from the INSEAD-Wharton Alliance, the Dreyfus Foundation and INSEAD alumnus Yves Burrus. It brought together leading scholars from a range of disciplines, institutions and journals for discussions on data management, open science, theory building and the growing role of AI in research integrity.

No scientific truth without integrity 

“As social scientists, our role is to reveal the truths that might explain and advance human society,” said Smith, ESRI’s Academic Director. In an era where other truth-seeking professions, from journalism to the judiciary, have been weakened by social media, political pressure and cost-cutting, academic research remains one of the few comparatively robust sources of truth.

“Integrity is the most fundamental of our professional values, alongside curiosity, objectivity, rigour and scepticism,” continued Smith. “Cases of academic fraud are dangerous because they shake public faith. Social scientists lose credibility with the policy makers, practitioners and populations they are supposed to serve. Research institutions, journals and conferences are damaged. The world looks to other – less reliable, more biased – sources for its truth.”

Stories like Fredrickson’s often start with good intentions, observed Leif Nelson of Berkeley Haas School of Business in his keynote speech. Nelson is the co-founder of Data Colada, which has exposed numerous cases of academic fraud. He is credited with helping to coin the term “p-hacking”, the practice of cherry-picking data to produce statistically significant results. Sometimes, he said, researchers want to find something so badly that they convince themselves they have. 

Spotting the problem doesn’t require a statistics degree. Most of Data Colada’s investigations have come from careful reading rather than technical analysis. Typical clues include data that looks too clean to be real, results that defy common sense, oddly formatted tables or poorly constructed survey questions. 

These red flags have been common in psychology, a field still grappling with a decade-long struggle to reproduce many of its best-known findings, now referred to as psychology’s replication crisis”. The effects have been felt acutely in business schools, where psychology underpins much of the research in organisational behaviour, marketing, decision sciences and behavioural economics. 

P is for p-hacking and psychology

Simine Vazire from the University of Melbourne, another keynote speaker, has built her career on reforming psychology’s research culture, and once crowdfunded a legal defence for the Data Colada team after a discredited Harvard Business School academic attempted to sue them for US$25 million.

The lessons from psychology’s reckoning aren’t always obvious, Vazire said in her address. Transparency, sharing all data and code from a study, is necessary but not sufficient on its own. Even with full cooperation from original authors, only 30% to 40% of conclusions in psychology papers can typically be replicated. 

One tool that has helped is preregistration, the practice of recording a study’s hypotheses and methods before it begins, which guards against bias and selective reporting after the fact. 

But Vazire argued that tools alone won’t fix the underlying culture. Researchers need to embrace scepticism of their own work, and journal publishers need to invest meaningfully in quality control rather than relying largely on unpaid peer reviewers.

The coming tsunami

Then there is the other kind of AI. One workshop panellist, a journal editor, reported a 42% increase in submissions since 2021/22, the year before ChatGPT’s release, alongside a marked decline in the quality of both papers and reviews. “The tsunami is coming,” warned another contributor, who observed that career incentives in academia have always provided a motive for cutting corners. But generative AI now supplies new means and opportunity – from fabricated data to sloppy reviews and plagiarism.

Yet technology also offers positive opportunities for research integrity. Algorithmic AI, distinct from its hallucination-prone generative cousin, relies on explicit programming and pre-defined rules to complete specific tasks. It is already proving useful for catching errors, checking equations and stripping data down to its underlying logic. Rather than banning AI outright from journal and conference submissions, panellists argued for policies that support its responsible use, with transparency about when and how it is used, and human judgement serving as the ultimate gatekeeper. 

What 2.9103 really teaches us

Whatever direction AI takes, there will always be people willing to commit fraud and people willing to believe them. But Strohminger sees a further lesson in the story of 2.9103. “Precision,” she warned, “sometimes makes models less accurate, not more.” Researchers should be wary of borrowing the trappings of hard sciences without the substance behind them. 

It might be tempting to treat cases like this as rare exceptions, best left unexamined for fear of casting doubt on the wider profession. The judges who dismissed the defamation claim against the Data Colada bloggers disagreed. In their view, calling out past mistakes is a fundamental part of how research is meant to work, not an act of malice. High-profile failures of integrity are damaging. But not exposing them would be worse.

The Workshop on Research Integrity was ESRI’s 2026 signature event. Visit the event page to view selected presentations from the keynote speakers and recommendations from the panel discussions.

Edited by:

Verity Ashton

About the author(s)

View Comments
No comments yet.
Leave a Comment
Please log in or sign up to comment.