NEWS
NEWS
Gender Bias in Generative AI: How Yesterday’s Visibility Gap Morphs into Tomorrow’s Knowledge Gap
Gender Bias in Generative AI: How Yesterday’s Visibility Gap Morphs into Tomorrow’s Knowledge Gap
Today, you can ask a generative AI system to name five experts on a specific topic, and the answers will be generated immediately and sound compelling. What you will never see, though, is the long list of those who never made it into that response.
While traditional search at least exposes competing pages, names, sources and other information, a synthesised answer compresses those options into just one response. This means that historical gaps in who has been published, cited and recognised can influence who appears authoritative to today’s LLMs.
Gender bias makes this particularly consequential. AI systems can reinforce an existing imbalance by repeatedly surfacing those whose expertise is already easier to find and verify. Put differently, this can shape who is perceived as an expert today, while those who were historically less visible risk remaining even more invisible down the road.
Gender Bias in Generative AI Starts Before the Answer Is Generated
In fact, gender bias in generative AI starts long before an LLM gives an output. AI systems cannot operate independently of the information environments around them, which means that their responses can be affected by training data, retrieval systems, ranking methods, design choices and publicly available web content.
So, if historical sources contain fewer women in visible positions, fewer profiles of their work or fewer references connecting their names to specific fields, an AI system may encounter an uneven evidence base before it generates any recommendation.
A 2024 UNESCO study on gender bias in large language models shows this quite well. It found that several LLMs associated women more frequently with domestic roles and family-related language, while men were more often linked to careers, business and higher-status occupations. Even though the study does not examine expert recommendations specifically, it perfectly displays how gendered patterns can reappear in model outputs.
Representation Bias Is Not the Same as Citation Disparity
If the UNESCO findings point specifically to representation bias, when it comes to expert visibility, though, there is a citation gap. These are two closely connected concepts, but they are not interchangeable.
Representation bias is about who is present in the information environment in the first place. Citation gaps, meanwhile, concern whose work is referenced, acknowledged and connected to a particular field of expertise.
This means that a woman may appear in articles, research databases or other public sources, yet still have fewer citations, references or explicit connections between her name and a specific subject area.
As a result, representation alone tells us very little about whether the available signals are sufficient for an AI system to recognise, retrieve and eventually recommend that person when a user asks for an expert. There must be enough public evidence for the system to connect that person consistently with a particular field of expertise.
Who Does AI Consider an Expert?
AI tends to consider someone an expert when there are strong and repeated indications connecting that person to a specific field, for example, publications, citations, references and other visible evidence of their work.
But who actually appears when an LLM is asked to name those experts? Researchers at the Complexity Science Hub tested this by asking six language models to recommend leading physicists, using a database of more than 450,000 scientists as a reference point.
Women accounted for around 14–32% of researchers in that database, depending on the subfield and period. Yet most of the tested LLMs recommended an even smaller proportion of women, and in some cases generated expert lists with no women at all. The models also tended to favour senior, highly cited male researchers from the US.
So, even where women are already present in the underlying expert pool, AI-generated recommendations can change this representation.
How To Reduce Gender Bias in AI Expert Recommendations
Reducing gender bias in AI expert recommendations requires work on two levels: the information environment from which expertise is discovered, and the systems that turn that information into recommendations.
Build a stronger public record of expertise
The first challenge is to verify how consistently someone’s name is connected to identifiable expertise. Original research, authored analysis, named commentary, interviews and expert contributions all create public evidence that links a person to specific ideas and subject areas.
This is also relevant to the work we do at Drofa Comms through Women Leading the Way. In finance, fintech and Web3, we regularly meet women whose expertise is already substantial, but whose public record does not always reflect the depth of their experience. Giving that knowledge an attributable space matters because it allows journalists, researchers, search systems and AI tools to encounter the expertise under the name of the person who actually holds it.
But this cannot become another responsibility placed mainly on women themselves. Media outlets decide whom they quote. Companies decide which executives receive bylines and speaking opportunities. Conference organisers decide whose names repeatedly appear on panels. These choices collectively influence whose expertise becomes easier to discover and verify over time.
Test the systems that recommend experts
The second part of the problem comes down to AI developers and the organisations deploying these systems. If a model repeatedly recommends the same demographic groups as experts, this regularity should itself be treated as something worth testing.
Developers can compare recommendations against relevant expert pools, examine whether retrieval methods change representation, and monitor whether improvements in factual accuracy come at the expense of diversity in the people being surfaced.
From our perspective, the goal should not be visibility for visibility’s sake. It is to make real expertise easier to find, attribute and verify, while also ensuring that AI systems are tested for the biases they may reproduce from the information environments they rely on.
The New Visibility Gap Could Be Harder to Spot
Traditional search at least makes its alternatives visible. Several pages can appear alongside one another, different experts can be quoted, and users can continue searching when the first result is not enough.
Generative AI changes this. It can turn multiple sources and signals into one seemingly complete answer, making it much harder to see who was considered, who was left out and why.
This is what makes the next visibility gap potentially more difficult to recognise. The risk is both that some experts may appear less often in AI-generated answers and that their absence can become almost invisible to the person asking the question.
If historical differences in visibility continue to influence who AI systems recognise and recommend, inequality may stop looking like exclusion. It may simply look like the answer.
Today, you can ask a generative AI system to name five experts on a specific topic, and the answers will be generated immediately and sound compelling. What you will never see, though, is the long list of those who never made it into that response.
While traditional search at least exposes competing pages, names, sources and other information, a synthesised answer compresses those options into just one response. This means that historical gaps in who has been published, cited and recognised can influence who appears authoritative to today’s LLMs.
Gender bias makes this particularly consequential. AI systems can reinforce an existing imbalance by repeatedly surfacing those whose expertise is already easier to find and verify. Put differently, this can shape who is perceived as an expert today, while those who were historically less visible risk remaining even more invisible down the road.
Gender Bias in Generative AI Starts Before the Answer Is Generated
In fact, gender bias in generative AI starts long before an LLM gives an output. AI systems cannot operate independently of the information environments around them, which means that their responses can be affected by training data, retrieval systems, ranking methods, design choices and publicly available web content.
So, if historical sources contain fewer women in visible positions, fewer profiles of their work or fewer references connecting their names to specific fields, an AI system may encounter an uneven evidence base before it generates any recommendation.
A 2024 UNESCO study on gender bias in large language models shows this quite well. It found that several LLMs associated women more frequently with domestic roles and family-related language, while men were more often linked to careers, business and higher-status occupations. Even though the study does not examine expert recommendations specifically, it perfectly displays how gendered patterns can reappear in model outputs.
Representation Bias Is Not the Same as Citation Disparity
If the UNESCO findings point specifically to representation bias, when it comes to expert visibility, though, there is a citation gap. These are two closely connected concepts, but they are not interchangeable.
Representation bias is about who is present in the information environment in the first place. Citation gaps, meanwhile, concern whose work is referenced, acknowledged and connected to a particular field of expertise.
This means that a woman may appear in articles, research databases or other public sources, yet still have fewer citations, references or explicit connections between her name and a specific subject area.
As a result, representation alone tells us very little about whether the available signals are sufficient for an AI system to recognise, retrieve and eventually recommend that person when a user asks for an expert. There must be enough public evidence for the system to connect that person consistently with a particular field of expertise.
Who Does AI Consider an Expert?
AI tends to consider someone an expert when there are strong and repeated indications connecting that person to a specific field, for example, publications, citations, references and other visible evidence of their work.
But who actually appears when an LLM is asked to name those experts? Researchers at the Complexity Science Hub tested this by asking six language models to recommend leading physicists, using a database of more than 450,000 scientists as a reference point.
Women accounted for around 14–32% of researchers in that database, depending on the subfield and period. Yet most of the tested LLMs recommended an even smaller proportion of women, and in some cases generated expert lists with no women at all. The models also tended to favour senior, highly cited male researchers from the US.
So, even where women are already present in the underlying expert pool, AI-generated recommendations can change this representation.
How To Reduce Gender Bias in AI Expert Recommendations
Reducing gender bias in AI expert recommendations requires work on two levels: the information environment from which expertise is discovered, and the systems that turn that information into recommendations.
Build a stronger public record of expertise
The first challenge is to verify how consistently someone’s name is connected to identifiable expertise. Original research, authored analysis, named commentary, interviews and expert contributions all create public evidence that links a person to specific ideas and subject areas.
This is also relevant to the work we do at Drofa Comms through Women Leading the Way. In finance, fintech and Web3, we regularly meet women whose expertise is already substantial, but whose public record does not always reflect the depth of their experience. Giving that knowledge an attributable space matters because it allows journalists, researchers, search systems and AI tools to encounter the expertise under the name of the person who actually holds it.
But this cannot become another responsibility placed mainly on women themselves. Media outlets decide whom they quote. Companies decide which executives receive bylines and speaking opportunities. Conference organisers decide whose names repeatedly appear on panels. These choices collectively influence whose expertise becomes easier to discover and verify over time.
Test the systems that recommend experts
The second part of the problem comes down to AI developers and the organisations deploying these systems. If a model repeatedly recommends the same demographic groups as experts, this regularity should itself be treated as something worth testing.
Developers can compare recommendations against relevant expert pools, examine whether retrieval methods change representation, and monitor whether improvements in factual accuracy come at the expense of diversity in the people being surfaced.
From our perspective, the goal should not be visibility for visibility’s sake. It is to make real expertise easier to find, attribute and verify, while also ensuring that AI systems are tested for the biases they may reproduce from the information environments they rely on.
The New Visibility Gap Could Be Harder to Spot
Traditional search at least makes its alternatives visible. Several pages can appear alongside one another, different experts can be quoted, and users can continue searching when the first result is not enough.
Generative AI changes this. It can turn multiple sources and signals into one seemingly complete answer, making it much harder to see who was considered, who was left out and why.
This is what makes the next visibility gap potentially more difficult to recognise. The risk is both that some experts may appear less often in AI-generated answers and that their absence can become almost invisible to the person asking the question.
If historical differences in visibility continue to influence who AI systems recognise and recommend, inequality may stop looking like exclusion. It may simply look like the answer.
London office
Rise, created by Barclays, 41 Luke St, London EC2A 4DP
Nicosia office
2043, Nikokreontos 29, office 202
DP FINANCE COMM LTD (#13523955) Registered Address: N1 7GU, 20-22 Wenlock Road, London, United Kingdom For Operations In The UK
AGAFIYA CONSULTING LTD (#HE 380737) Registered Address: 2043, Nikokreontos 29, Flat 202, Strovolos, Cyprus For Operations In The EU, LATAM, United Stated Of America And Provision Of Services Worldwide
Drofa © 2024
London office
Rise, created by Barclays, 41 Luke St, London EC2A 4DP
Nicosia office
2043, Nikokreontos 29, office 202
DP FINANCE COMM LTD (#13523955) Registered Address: N1 7GU, 20-22 Wenlock Road, London, United Kingdom For Operations In The UK
AGAFIYA CONSULTING LTD (#HE 380737) Registered Address: 2043, Nikokreontos 29, Flat 202, Strovolos, Cyprus For Operations In The EU, LATAM, United Stated Of America And Provision Of Services Worldwide
Drofa © 2024
