Jacob Metoyerresearch / computation / making

Independent research

Working with AI, carefully.

A hobby that got serious: I build systems where AI models do real work, and then check whether they actually did it well.

Research

Independent research — ongoing

Independent AI systems research

This started as a hobby and turned into something I run like research. I build systems where one or several AI models work on real tasks, then look closely at what happened: what they got right, where they failed, and whether the result was worth the extra coordination.

The habit I care about most is keeping the evidence. Tasks, tool actions, outputs and failures get saved, so a run can be checked later instead of remembered generously. I'm especially interested in when a piece of reasoning that worked once can become a procedure that works reliably.

It's independent work, done outside coursework and outside any lab, and none of it has been through peer review.

Independent work, done outside coursework and outside any lab, and not peer reviewed.

The questions

Does running more models in parallel help, or does it mostly add coordination overhead? How much context does a task really need? And when a model gets something right once, can that turn into a procedure you can rely on, or was it luck?

The easy trap is to confuse fast with good. So I try to keep the comparison fair: the same task done by one model, by a small group, and by a larger setup, with the failures kept alongside the successes.

Where it could go

My own work is the first test bed. Eventually I’d like to try this with organizations: look at how a team already does something, build a small AI-assisted version with ordinary tools like ChatGPT, and measure whether it’s better on quality, time and cost. I haven’t done that for a client yet. It’s where I want this to go.