Skip to content

Latest commit

 

History

19 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 

Repository files navigation

Keeping AI From Going Rogue

Landing page for my AI safety work. Live at https://maricairns.github.io/build/

Helping AI Models Steer Clear of Shady Stuff

My career focus is to apply psychology and human factors expertise to helping make AI go well. I believe human factors domain expertise is vital at the source of the river. Where AI models are built. Where AI is trained, fine-tuned, aligned, evaluated, blue-teamed, red-teamed (and black-teamed), and stress-tested.

Especially when it comes to emerging misalignment, alignment faking, and hidden objectives in AI models. In particular, where AI models display behaviours that align with dark triad characteristics and personality traits. To help account for these AI model behaviours before deployment — and the more covert, dormant-over-time, hidden objectives that may only manifest after deployment, with time-lags.

As AI systems scale and become increasingly powerful, getting this right becomes existential.

Focus

  • AI Model Behaviour
  • AI Model Personality
  • AI Model (meta)Cognition

The page

Two routes in: Work with Me and Fund this Work.


Dr Mari Cairns — DClinPsych (University of Oxford), CPsychol, AFBPsS, MBA

people · processes · data · technology

About

Source for my site on AI model and AI agent behaviour, manipulation, deception, personality, and persona vectors.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages