Allium User Group 2026-08-05

Here’s a link to the recording: you will need to be logged into Zoom to view it.

Meeting Summary

Testing across models, and early thinking on Allium v4

This week’s call centred on how we measure whether changes to Allium make things better or worse. A test rig has been in use that runs the current version alongside a modified one, producing a score for what worked, how many tokens were used and where the results differ. It builds on the idea that evals and tests do different jobs. Some things you can check outright, like whether the spec is correct, while others you can only get a feel for. The same setup could be pointed at different models, which is useful given that not everyone trusts Allium to work well with, say, Gemini.

The other main topic was where v4 might go. After a look at prior art, from TLA+ to formal proof tools, the aim is to hit the sweet spot from design through to implementation, which existing tools miss. One idea we’re considering doubling down on is pulling in specs as libraries, so you could rely on a ‘spec for Kafka’ or a piece of infrastructure rather than explaining it all to the AI. This raises real trade-offs and may bring breaking changes.

@yavor.panayotov has used Allium to turn JDK docs into 10 specs and 645 obligations, finding 15 inconsistencies in the process. The article describing the approach in detail is available here.


If you’d like to attend any future meetings, (currently Wednesdays, 3-4pm BST) please register your interest here. There’s no obligation to turn up once registered, it just lets us know who to include on administrative notifications. The session lasts for an hour, but you’re welcome to drop in and out any time.

In the future we may host in-person London meet-ups as well, do let us know if you have any interest in attending one.