Autoformalization in VacuumWorld
Nine-constraint symbolic evaluator · F1-vs-N sweep
Percepts to Prolog knowledge bases, with an evaluation harness built to tell induced rules apart from enumerated facts.
Open to research collaborations & Summer 2027 internships
Undergraduate Research Fellow · DICE Lab, Royal Holloway
I work on mechanistic interpretability and neurosymbolic AI. My current question: when a language model writes a correct formal specification, has it induced a general rule, or has it enumerated the cases it happened to see?
Current work
Autoformalization at RHUL's DICE Lab, supervised by Prof. Kostas Stathis and Dr. Agnieszka Mensfelt.
A model that writes wall(0, Y) for every Y it has seen looks
exactly like a model that has learned what a wall is — until you make the
grid bigger.
I translate VacuumWorld agent percepts into Prolog knowledge bases and ask whether the model induced a general boundary-wall rule or hardcoded enumerated facts. Scoring is done by a nine-constraint symbolic evaluator rather than string matching, and an F1-vs-N grid-size sweep separates real rule induction from enumeration that only appears correct at small N. Nine models were run across four prompt conditions, with a composition table isolating each scaffolding factor as a single-variable contrast.
Selected work
Nine-constraint symbolic evaluator · F1-vs-N sweep
Percepts to Prolog knowledge bases, with an evaluation harness built to tell induced rules apart from enumerated facts.
Users across 18+ countries
Burmese literacy platform on Cloudflare Workers AI — a multilingual NLP pipeline served from the edge.
Voice to settled payment, under five seconds
Cross-border payments prototype built at the LSE Agentic Society Hackathon: speech in, on-chain settlement out.
Writing
What I am looking for
Interpretability, neurosymbolic reasoning, and evaluation. I am most useful where someone needs an experiment designed carefully and a harness that will not flatter the results.