Landing page for my AI safety work. Live at https://maricairns.github.io/build/
Helping AI Models Steer Clear of Shady Stuff
My career focus is to apply psychology and human factors expertise to helping make AI go well. I believe human factors domain expertise is vital at the source of the river. Where AI models are built. Where AI is trained, fine-tuned, aligned, evaluated, blue-teamed, red-teamed (and black-teamed), and stress-tested.
Especially when it comes to emerging misalignment, alignment faking, and hidden objectives in AI models. In particular, where AI models display behaviours that align with dark triad characteristics and personality traits. To help account for these AI model behaviours before deployment — and the more covert, dormant-over-time, hidden objectives that may only manifest after deployment, with time-lags.
As AI systems scale and become increasingly powerful, getting this right becomes existential.
- AI Model Behaviour
- AI Model Personality
- AI Model (meta)Cognition
Two routes in: Work with Me and Fund this Work.
Dr Mari Cairns — DClinPsych (University of Oxford), CPsychol, AFBPsS, MBA
people · processes · data · technology