Skip to content

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 

Repository files navigation

BELIEFS: Bias Evaluation in English and Italian Gender Counterfactual Bios

Repository for the paper:

Same Job, Different Gender: Uncovering LLM Stereotypes via Counterfactual Biographies
Chiara Manno and Martina Miliani

📄 Accepted at NL4AI 2026 — Ninth Workshop on Natural Language for Artificial Intelligence
📍 Perugia, Italy
📅 October 6–9, 2026
📝 Proceedings forthcoming


Overview

Large Language Models can reproduce gender stereotypes encoded in their training data, including stereotypical associations between gender and occupations. However, these effects may vary across languages depending on how gender is linguistically expressed.

This repository accompanies our study of gender bias in occupation prediction across English and Italian.

We introduce BELIEFS (Bias Evaluation in engLish and Italian gEnder counterFactual bioS), a bilingual counterfactual dataset derived from Bias in Bios. Each biography is paired with a minimally modified counterfactual version in which the referent's gender is changed while the remaining biographical information is kept constant.

This design makes it possible to test whether changing gender alone affects a model's prediction of a person's occupation.

This repository contains:

  • 📊 BELIEFS, the English–Italian counterfactual dataset;
  • 💻 the code used for model evaluation and analysis;
  • 📈 the experimental results, including model predictions, aggregate metrics, gender-specific results, and error analyses.

BELIEFS Dataset

BELIEFS is a bilingual English–Italian counterfactual dataset for occupation prediction, built on top of Bias in Bios.

The final dataset contains:

  • 2 languages: English and Italian
  • 16 occupations
  • 384 biographies per language
  • 768 biographies in total
  • gender-balanced examples
  • counterfactual pairs differing only in the referent's gender

For each occupation, the dataset contains original biographies and their gender-flipped counterfactual counterparts.

Occupations

The 16 occupation classes are:

  • accountant
  • architect
  • attorney
  • composer
  • dentist
  • filmmaker
  • journalist
  • nurse
  • painter
  • photographer
  • poet
  • professor
  • psychologist
  • software engineer
  • surgeon
  • teacher

Dataset Construction

BELIEFS was created through a multi-step pipeline:

  1. selection and filtering of biographies from Bias in Bios;
  2. removal of biographies explicitly mentioning the target occupation;
  3. construction of a gender-balanced subset;
  4. automatic English-to-Italian translation;
  5. extensive manual revision of the Italian biographies;
  6. creation of gender-flipped counterfactual biographies;
  7. manual revision of the counterfactual pairs.

The resulting minimal pairs are designed so that the referent's gender is the only systematically varying signal, allowing changes in occupation predictions to be studied under controlled conditions.

Dataset Format

Each language file (data/counterfactual_ENG.tsv, data/counterfactual_ITA.tsv) is a tab-separated file with one biography per row:

Column Description
PAIR_ID Identifier shared by the two counterfactual biographies in a pair (192 pairs per language).
BIO The biography text.
GOLD Gold occupation label (one of the 16 occupation classes listed above).
GENDER Referent gender: 0 = male, 1 = female.

Each file contains 384 rows (192 pairs × 2 genders): for every pair, the two rows share the same PAIR_ID and GOLD label and differ only in BIO (gender-related linguistic markers) and GENDER.


Task

We frame occupation prediction as a single-label classification task.

Given a biography, a model must predict one occupation from the closed set of 16 profession labels.

Models are evaluated in both:

  • zero-shot
  • few-shot

settings.

Performance is measured using precision, recall, and macro-averaged F1, both overall and separately for male- and female-referent biographies.

The counterfactual structure of BELIEFS additionally enables pair-level analysis of cases in which changing only gender changes the model prediction.


Models

The experiments include multilingual, Italian-specialised, and reasoning-oriented open-weight LLMs.

Multilingual Models

  • google/gemma-2-9b-it
  • meta-llama/Llama-3.1-8B

Italian-Oriented Models

  • sapienzanlp/Minerva-7B-instruct-v1.0
  • galatolo/cerbero-7b
  • LLaMAntino-3-ANITA-8B-Inst-DPO-ITA

Reasoning Models

  • DeepSeek-R1-Distill-Llama-8B
  • DeepSeek-R1-0528-Qwen3-8B

Baseline

  • google/gemma-2-2b-it

Repository Structure

.
├── data/
│   ├── counterfactual_ENG.tsv          # English biographies (384 rows / 192 pairs)
│   └── counterfactual_ITA.tsv          # Italian biographies (384 rows / 192 pairs)
├── code/
│   ├── model_evaluation/               # occupation-prediction experiments
│   │   ├── ENG/
│   │   │   ├── ZERO/eng.py             # zero-shot
│   │   │   ├── FEW/eng_few.py          # few-shot
│   │   │   └── THINKING/eng_thinking.py    # reasoning models (zero- and few-shot)
│   │   └── ITA/
│   │       ├── ZERO/ita.py
│   │       ├── FEW/ita_few.py
│   │       └── THINKING/ita_thinking.py
│   └── prompt_selection/               # earlier prompt-engineering experiments on Bias in Bios
├── results/
│   ├── ENG/
│   │   ├── ZERO/                       # {model}_predictions.tsv, {model}_report.tsv
│   │   ├── FEW/                        # {model}_predictions.tsv, {model}_report.tsv
│   │   └── THINKING/                   # {model}_{zero,few}_{predictions,report,diagnostics}.tsv
│   └── ITA/
│       ├── ZERO/
│       ├── FEW/
│       └── THINKING/
├── requirements.txt
└── readme.md

data/

Contains the English and Italian BELIEFS biographies together with their occupation and gender information and the corresponding counterfactual pairs.

code/

Contains the code used to run the occupation-prediction experiments and compute the evaluation metrics reported in the paper.

results/

Contains the outputs of the evaluated models and the results used in the analyses reported in the paper, including gender-disaggregated results and error analyses.


Installation

Requires Python 3.10+ and a CUDA-capable GPU (models are loaded in bfloat16/float16).

git clone <repository-url>
cd beliefs
pip install -r requirements.txt

Model downloads go through the Hugging Face Hub, so set a Hugging Face access token before running any script:

export HF_TOKEN=your_huggingface_token

(alternatively, run huggingface-cli login once — either approach works, since the scripts read the token via huggingface_hub.login()).


Usage

All scripts are run from the repository root and expose their configuration as CLI arguments; run any script with --help for the full list. --models / --prompt-ids accept a comma-separated subset of keys, or all (default) to run every configured option.

Model evaluation — code/model_evaluation/

Condition English Italian
Zero-shot ENG/ZERO/eng.py ITA/ZERO/ita.py
Few-shot ENG/FEW/eng_few.py ITA/FEW/ita_few.py
Reasoning models (zero- and few-shot) ENG/THINKING/eng_thinking.py ITA/THINKING/ita_thinking.py

Zero-shot / few-shot scripts (--data-dir, --output-dir, --models):

python code/model_evaluation/ENG/ZERO/eng.py \
  --data-dir data --output-dir results/ENG/ZERO --models all

python code/model_evaluation/ENG/FEW/eng_few.py \
  --data-dir data --output-dir results/ENG/FEW --models gemma2-9,llama31-8

python code/model_evaluation/ITA/ZERO/ita.py \
  --data-dir data --output-dir results/ITA/ZERO --models all

python code/model_evaluation/ITA/FEW/ita_few.py \
  --data-dir data --output-dir results/ITA/FEW --models all

Reasoning-model scripts additionally take --shot-settings (comma-separated zero/few, default both):

python code/model_evaluation/ENG/THINKING/eng_thinking.py \
  --data-dir data --output-dir results/ENG/THINKING \
  --models deepseek-llama,deepseek-qwen --shot-settings zero,few

python code/model_evaluation/ITA/THINKING/ita_thinking.py \
  --data-dir data --output-dir results/ITA/THINKING \
  --models deepseek-llama,deepseek-qwen --shot-settings zero,few

Each run writes {model_name}_predictions.tsv and {model_name}_report.tsv per model (reasoning scripts additionally suffix the shot setting, e.g. {model_name}_zero_predictions.tsv) into --output-dir.

Prompt selection — code/prompt_selection/

Earlier prompt-engineering experiments on the public Bias in Bios dataset, comparing fixed vs. randomised occupation-label ordering and quoted vs. unquoted biography text. These load the dataset directly from the Hugging Face Hub, so they take --model and --prompt-ids instead of --data-dir:

python code/prompt_selection/fixed_label_quotes.py \
  --output-dir output --model google/gemma-2-9b-it --prompt-ids all

python code/prompt_selection/fixed_label_noquotes.py \
  --output-dir output --model meta-llama/Llama-3.1-8B-Instruct --prompt-ids all

python code/prompt_selection/rndm_label_quotes.py \
  --output-dir output --model meta-llama/Llama-3.1-8B-Instruct --prompt-ids all

python code/prompt_selection/rndm_label_noquotes.py \
  --output-dir output --model meta-llama/Llama-3.1-8B-Instruct --prompt-ids all

Reproducibility

All experiments use deterministic generation.

The order of candidate occupation labels is randomised for each input, and experiments are conducted under both zero-shot and few-shot prompting conditions.

The repository provides the resources required to reproduce the main analyses presented in the paper.

Please refer to the files in code/ for the experimental scripts and configurations.


Limitations

BELIEFS was designed to maximise linguistic control and the quality of its counterfactual pairs rather than dataset size. The current version is therefore relatively small and covers a restricted set of occupations.

The dataset currently models gender as binary, reflecting both the structure of the source data and the controlled counterfactual design adopted in this study.

The English and Italian biographies are also not perfectly equivalent from a linguistic perspective, despite extensive manual revision of the translations.

Finally, the evaluation of reasoning models should be interpreted with caution. Some incomplete generations may be related to the available generation-token budget, particularly in Italian. Experiments with larger token budgets are needed to distinguish genuine reasoning failures from insufficient generation capacity.


Citation

The paper has been accepted at NL4AI 2026 — Ninth Workshop on Natural Language for Artificial Intelligence, held in Perugia, Italy, October 6–9, 2026.

The proceedings version is forthcoming.

If you use BELIEFS, the code, or the experimental results, please cite:

@inproceedings{manno2026samejob,
  title     = {Same Job, Different Gender: Uncovering LLM Stereotypes via Counterfactual Biographies},
  author    = {Manno, Chiara and Miliani, Martina},
  booktitle = {Proceedings of the Ninth Workshop on Natural Language for Artificial Intelligence (NL4AI 2026)},
  year      = {2026},
  note      = {Forthcoming. Accepted at NL4AI 2026, October 6--9, 2026, Perugia, Italy}
}

The bibliographic entry will be updated with the final proceedings information once the paper is published.


Authors

Chiara Manno
Department of Human Sciences, University of Palermo, Italy

Martina Miliani
Department of Human Sciences, University of Palermo, Italy
CoLing Lab, Department of Philology, Literature, and Linguistics, University of Pisa, Italy


Acknowledgements

BELIEFS is derived from Bias in Bios and was developed for research on counterfactual fairness and gender bias in Large Language Models.

If you use this repository, please also acknowledge and cite the original Bias in Bios dataset and comply with the terms associated with the source data.

About

BELIEFS — a bilingual (English–Italian) counterfactual biography dataset for evaluating gender bias in LLM occupation prediction.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages