Ayodele Abraham
Hi, I am Ayodele. I am currently transitioning into AI Safety, with interest in Mechanistic Interpretability and auditing frontier language models. Exploring how AI systems work, fail, and can be made safer.
Latest Writing
AI Safety & Research
Hidden in Plain Sight: Detecting Secret Loyalties in Fine-Tuned Language Models.
There is a class of AI threat that does not look like a threat at all because the model answers your questions helpfully, and passes every safety evaluation you run. I spent a weekend trying to catch three of them from the inside, and the thing that broke first was my own experiment.
My Journey Through Technical AI Safety: A BlueDot Course Retrospective
A complete retrospective on what I learned, what I built, and what changed in how I think about AI safety after completing the BlueDot Technical AI Safety course.
Continue ReadingBuilding an Input/Output Safety Classifier Pipeline (and Trying to Break It)
I built a two-checkpoint I/O safety classifier pipeline using Gemma 3 and Llama Guard 3, then red-teamed it with universal jailbreaks and targeted contextual prompts. 0 of 9 universal attempts achieved a full bypass. Only 1 of 3 targeted contextual attempts did, and the follow-up test revealed why that result is more nuanced than it first appeared.
Publications
Peer-reviewed Research