Repository for the paper:
Same Job, Different Gender: Uncovering LLM Stereotypes via Counterfactual Biographies
Chiara Manno and Martina Miliani
📄 Accepted at NL4AI 2026 — Ninth Workshop on Natural Language for Artificial Intelligence
📍 Perugia, Italy
📅 October 6–9, 2026
📝 Proceedings forthcoming
Large Language Models can reproduce gender stereotypes encoded in their training data, including stereotypical associations between gender and occupations. However, these effects may vary across languages depending on how gender is linguistically expressed.
This repository accompanies our study of gender bias in occupation prediction across English and Italian.
We introduce BELIEFS (Bias Evaluation in engLish and Italian gEnder counterFactual bioS), a bilingual counterfactual dataset derived from Bias in Bios. Each biography is paired with a minimally modified counterfactual version in which the referent's gender is changed while the remaining biographical information is kept constant.
This design makes it possible to test whether changing gender alone affects a model's prediction of a person's occupation.
This repository contains:
- 📊 BELIEFS, the English–Italian counterfactual dataset;
- 💻 the code used for model evaluation and analysis;
- 📈 the experimental results, including model predictions, aggregate metrics, gender-specific results, and error analyses.
BELIEFS is a bilingual English–Italian counterfactual dataset for occupation prediction, built on top of Bias in Bios.
The final dataset contains:
- 2 languages: English and Italian
- 16 occupations
- 384 biographies per language
- 768 biographies in total
- gender-balanced examples
- counterfactual pairs differing only in the referent's gender
For each occupation, the dataset contains original biographies and their gender-flipped counterfactual counterparts.
The 16 occupation classes are:
- accountant
- architect
- attorney
- composer
- dentist
- filmmaker
- journalist
- nurse
- painter
- photographer
- poet
- professor
- psychologist
- software engineer
- surgeon
- teacher
BELIEFS was created through a multi-step pipeline:
- selection and filtering of biographies from Bias in Bios;
- removal of biographies explicitly mentioning the target occupation;
- construction of a gender-balanced subset;
- automatic English-to-Italian translation;
- extensive manual revision of the Italian biographies;
- creation of gender-flipped counterfactual biographies;
- manual revision of the counterfactual pairs.
The resulting minimal pairs are designed so that the referent's gender is the only systematically varying signal, allowing changes in occupation predictions to be studied under controlled conditions.
Each language file (data/counterfactual_ENG.tsv, data/counterfactual_ITA.tsv) is a tab-separated file with one biography per row:
| Column | Description |
|---|---|
PAIR_ID |
Identifier shared by the two counterfactual biographies in a pair (192 pairs per language). |
BIO |
The biography text. |
GOLD |
Gold occupation label (one of the 16 occupation classes listed above). |
GENDER |
Referent gender: 0 = male, 1 = female. |
Each file contains 384 rows (192 pairs × 2 genders): for every pair, the two rows share the same PAIR_ID and GOLD label and differ only in BIO (gender-related linguistic markers) and GENDER.
We frame occupation prediction as a single-label classification task.
Given a biography, a model must predict one occupation from the closed set of 16 profession labels.
Models are evaluated in both:
- zero-shot
- few-shot
settings.
Performance is measured using precision, recall, and macro-averaged F1, both overall and separately for male- and female-referent biographies.
The counterfactual structure of BELIEFS additionally enables pair-level analysis of cases in which changing only gender changes the model prediction.
The experiments include multilingual, Italian-specialised, and reasoning-oriented open-weight LLMs.
google/gemma-2-9b-itmeta-llama/Llama-3.1-8B
sapienzanlp/Minerva-7B-instruct-v1.0galatolo/cerbero-7bLLaMAntino-3-ANITA-8B-Inst-DPO-ITA
DeepSeek-R1-Distill-Llama-8BDeepSeek-R1-0528-Qwen3-8B
google/gemma-2-2b-it
.
├── data/
│ ├── counterfactual_ENG.tsv # English biographies (384 rows / 192 pairs)
│ └── counterfactual_ITA.tsv # Italian biographies (384 rows / 192 pairs)
├── code/
│ ├── model_evaluation/ # occupation-prediction experiments
│ │ ├── ENG/
│ │ │ ├── ZERO/eng.py # zero-shot
│ │ │ ├── FEW/eng_few.py # few-shot
│ │ │ └── THINKING/eng_thinking.py # reasoning models (zero- and few-shot)
│ │ └── ITA/
│ │ ├── ZERO/ita.py
│ │ ├── FEW/ita_few.py
│ │ └── THINKING/ita_thinking.py
│ └── prompt_selection/ # earlier prompt-engineering experiments on Bias in Bios
├── results/
│ ├── ENG/
│ │ ├── ZERO/ # {model}_predictions.tsv, {model}_report.tsv
│ │ ├── FEW/ # {model}_predictions.tsv, {model}_report.tsv
│ │ └── THINKING/ # {model}_{zero,few}_{predictions,report,diagnostics}.tsv
│ └── ITA/
│ ├── ZERO/
│ ├── FEW/
│ └── THINKING/
├── requirements.txt
└── readme.md
Contains the English and Italian BELIEFS biographies together with their occupation and gender information and the corresponding counterfactual pairs.
Contains the code used to run the occupation-prediction experiments and compute the evaluation metrics reported in the paper.
Contains the outputs of the evaluated models and the results used in the analyses reported in the paper, including gender-disaggregated results and error analyses.
Requires Python 3.10+ and a CUDA-capable GPU (models are loaded in bfloat16/float16).
git clone <repository-url>
cd beliefs
pip install -r requirements.txtModel downloads go through the Hugging Face Hub, so set a Hugging Face access token before running any script:
export HF_TOKEN=your_huggingface_token(alternatively, run huggingface-cli login once — either approach works, since the scripts read the token via huggingface_hub.login()).
All scripts are run from the repository root and expose their configuration as CLI arguments; run any script with --help for the full list. --models / --prompt-ids accept a comma-separated subset of keys, or all (default) to run every configured option.
| Condition | English | Italian |
|---|---|---|
| Zero-shot | ENG/ZERO/eng.py |
ITA/ZERO/ita.py |
| Few-shot | ENG/FEW/eng_few.py |
ITA/FEW/ita_few.py |
| Reasoning models (zero- and few-shot) | ENG/THINKING/eng_thinking.py |
ITA/THINKING/ita_thinking.py |
Zero-shot / few-shot scripts (--data-dir, --output-dir, --models):
python code/model_evaluation/ENG/ZERO/eng.py \
--data-dir data --output-dir results/ENG/ZERO --models all
python code/model_evaluation/ENG/FEW/eng_few.py \
--data-dir data --output-dir results/ENG/FEW --models gemma2-9,llama31-8
python code/model_evaluation/ITA/ZERO/ita.py \
--data-dir data --output-dir results/ITA/ZERO --models all
python code/model_evaluation/ITA/FEW/ita_few.py \
--data-dir data --output-dir results/ITA/FEW --models allReasoning-model scripts additionally take --shot-settings (comma-separated zero/few, default both):
python code/model_evaluation/ENG/THINKING/eng_thinking.py \
--data-dir data --output-dir results/ENG/THINKING \
--models deepseek-llama,deepseek-qwen --shot-settings zero,few
python code/model_evaluation/ITA/THINKING/ita_thinking.py \
--data-dir data --output-dir results/ITA/THINKING \
--models deepseek-llama,deepseek-qwen --shot-settings zero,fewEach run writes {model_name}_predictions.tsv and {model_name}_report.tsv per model (reasoning scripts additionally suffix the shot setting, e.g. {model_name}_zero_predictions.tsv) into --output-dir.
Earlier prompt-engineering experiments on the public Bias in Bios dataset, comparing fixed vs. randomised occupation-label ordering and quoted vs. unquoted biography text. These load the dataset directly from the Hugging Face Hub, so they take --model and --prompt-ids instead of --data-dir:
python code/prompt_selection/fixed_label_quotes.py \
--output-dir output --model google/gemma-2-9b-it --prompt-ids all
python code/prompt_selection/fixed_label_noquotes.py \
--output-dir output --model meta-llama/Llama-3.1-8B-Instruct --prompt-ids all
python code/prompt_selection/rndm_label_quotes.py \
--output-dir output --model meta-llama/Llama-3.1-8B-Instruct --prompt-ids all
python code/prompt_selection/rndm_label_noquotes.py \
--output-dir output --model meta-llama/Llama-3.1-8B-Instruct --prompt-ids allAll experiments use deterministic generation.
The order of candidate occupation labels is randomised for each input, and experiments are conducted under both zero-shot and few-shot prompting conditions.
The repository provides the resources required to reproduce the main analyses presented in the paper.
Please refer to the files in code/ for the experimental scripts and configurations.
BELIEFS was designed to maximise linguistic control and the quality of its counterfactual pairs rather than dataset size. The current version is therefore relatively small and covers a restricted set of occupations.
The dataset currently models gender as binary, reflecting both the structure of the source data and the controlled counterfactual design adopted in this study.
The English and Italian biographies are also not perfectly equivalent from a linguistic perspective, despite extensive manual revision of the translations.
Finally, the evaluation of reasoning models should be interpreted with caution. Some incomplete generations may be related to the available generation-token budget, particularly in Italian. Experiments with larger token budgets are needed to distinguish genuine reasoning failures from insufficient generation capacity.
The paper has been accepted at NL4AI 2026 — Ninth Workshop on Natural Language for Artificial Intelligence, held in Perugia, Italy, October 6–9, 2026.
The proceedings version is forthcoming.
If you use BELIEFS, the code, or the experimental results, please cite:
@inproceedings{manno2026samejob,
title = {Same Job, Different Gender: Uncovering LLM Stereotypes via Counterfactual Biographies},
author = {Manno, Chiara and Miliani, Martina},
booktitle = {Proceedings of the Ninth Workshop on Natural Language for Artificial Intelligence (NL4AI 2026)},
year = {2026},
note = {Forthcoming. Accepted at NL4AI 2026, October 6--9, 2026, Perugia, Italy}
}The bibliographic entry will be updated with the final proceedings information once the paper is published.
Chiara Manno
Department of Human Sciences, University of Palermo, Italy
Martina Miliani
Department of Human Sciences, University of Palermo, Italy
CoLing Lab, Department of Philology, Literature, and Linguistics, University of Pisa, Italy
BELIEFS is derived from Bias in Bios and was developed for research on counterfactual fairness and gender bias in Large Language Models.
If you use this repository, please also acknowledge and cite the original Bias in Bios dataset and comply with the terms associated with the source data.