ayabraham.com
AYODELE
AI Safety · Mech. Interp.AI Auditing · Alignment
AI Safety
Mechanistic Interpretability
AI Auditing
AI Safety

Ayodele Abraham

Hi, I am Ayodele. I am currently transitioning into AI Safety, with interest in Mechanistic Interpretability and auditing frontier language models. Exploring how AI systems work, fail, and can be made safer.

View ResearchRead Writing

Latest Writing


AI Safety & Research
AI Safety

Hidden in Plain Sight: Detecting Secret Loyalties in Fine-Tuned Language Models.

There is a class of AI threat that does not look like a threat at all because the model answers your questions helpfully, and passes every safety evaluation you run. I spent a weekend trying to catch three of them from the inside, and the thing that broke first was my own experiment.

Read →
AI Safety

My Journey Through Technical AI Safety: A BlueDot Course Retrospective

A complete retrospective on what I learned, what I built, and what changed in how I think about AI safety after completing the BlueDot Technical AI Safety course.

Continue Reading
AI Safety

Building an Input/Output Safety Classifier Pipeline (and Trying to Break It)

I built a two-checkpoint I/O safety classifier pipeline using Gemma 3 and Llama Guard 3, then red-teamed it with universal jailbreaks and targeted contextual prompts. 0 of 9 universal attempts achieved a full bypass. Only 1 of 3 targeted contextual attempts did, and the follow-up test revealed why that result is more nuanced than it first appeared.

Read →
View all AI Safety writing →

Publications


Peer-reviewed Research
01.
Prediction of tool wear based on GA-BP neural network
Proceedings of the Institution of Mechanical Engineers, Part B: Journal of Engineering Manufacture (2022)
Sage Journals DOI →
02.
Surface roughness and chip morphology of wood-plastic composites manufactured via high-speed milling
BioResources, 16(3), 2021
BioResources →
View all publications →