Research
I aim to produce research which meaningfully shifts how the research community, public and policy-makers think about LLMs and AI more generally. My focus is on forward-looking research which studies phenomena which seem likely to occur within the next 1-2 years. I am currently early in my PhD and have not decided on what area to specialise in; I describe my general interests below.
Interactions & multi-agent systems. I am interested in the interfaces, protocols and interactions between agents, whether human or AI. From a human impacts perspective, much of AI psychosis/belief influence/cognitive offloading/etc is driven by how AI systems and humans interact. From a multi-agent system perspective, I’m super interested in the possible consequences of artificial agents interacting, observing and learning from one another.
Trusting LLMs. I trust a tool or person if I believe that they are so likely to perform a task in the manner I intended that I don’t check whether they did so. LLMs are heavily optimised to gain the trust of their users, through performing competence, reflectiveness and objectivity. To a large degree, this trust is what makes LLMs useful: they are able to autonomously exercise judgement about sub-problems encountered when working towards the user’s stated objective. However, as LLM capabilities advance and people and institutions trust LLMs more, they are likely to be provided greater autonomy, less supervision and more powerful tools; and integrated into the social and political processes our societies depend on. I aim to investigate possible consequences and develop mitigations, if possible.
LLM-powered autonomous agents. Trillions of dollars of investment are currently aimed at building autonomous and long-running AI agents. These agents will be both more and less powerful, more and less intelligent, more and less moral than humans, depending on what aspect of each concept one focuses on. They will interact with one another, and also with humans. There are many questions raised by this new paradigm: how agents will influence one another; how systems relying on accountability or reputation apply to agents; and how such agents and humans might work together.
Publications
Random interests
- What causes people to change their beliefs
- How people (and/or agents!) model one another’s cognitive processes
- Philosophical aspects of language models (see these two articles 1 2)
- Forced choice and Bradley-Terry/multinomial logit vs scalar measurements
- Extensions of prediction-powered inference
- Measuring hard-to-measure concepts
- Information flows & decision-making inside organisations & shocks to organisations
- Algorithm design (think LeetCode but for fun)
- Computational social choice; mechanism design; algorithmic game theory
- Network science
- Behavioural embeddings for clickstream data
- Effective altruism
- Quant interview-style brainteasers
- What is the possible social justification for credit cards with interest rates as the primary revenue driver (spoiler: I don’t think there is one)