If AI and code run in separate lanes, numerical errors are impossible.But which formula applies is itself a model judgment. Where's the deterministic boundary?
A Sphere Event Series
Ground Truth
ground truth, noun · In machine learning, the verified correct answer a model is graded against. An invite-only afternoon for engineers working at the intersection of AI x regulated industries.
The Premise for the Day
A half-day event for senior engineers interested in building where AI and regulated industries with zero room for error meet.
For most AI applications, ground truth ends at the eval set. Models ship on benchmark scores and A/B tests, and "good enough" is a product decision.
Compliance doesn't work that way. The jurisdiction sets the rate. The authority holds the position. The filing is right or it's wrong. Ground Truth brings together the senior engineers building to the bar where there is zero margin for error - tax, finance, healthcare, legal, payroll - for a half-day of technical and in-depth discussion.
What Ground Truth will cover
Two panels of senior engineers from companies like Anthropic, Harvey, and more, discussing real technical questions at the heart of building. No ambiguous keynotes, no sales decks. Then, afterwards, an evening to connect with peers interested in solving difficult problems: curated 1:1s and a long reception.
If an LLM judges an LLM and they share training lineage, are you measuring quality or shared bias?
A schema-valid answer can still be flat wrong.Where has structured output bought your team false confidence?
An agent step fails mid-plan and retries.What guarantees the card doesn't get charged twice?
Reviewers who reject almost nothing: are they reviewing, or signing?
When a human overrides the model, is that automatically ground truth?
Where does jurisdiction live in your stack: a deterministic gate, or a system prompt an injection can rewrite?
Is human review a moat you're deepening or a cost you're racing to eliminate?
Where does your domain advantage live: weights, ontology, evals, retrieval?And which part is worthless when the next model ships?
Years of audit trail can't be rebuilt at a competitor.Deliberate architecture, or lucky accident?
Is your data a compounding network effect?Or is it a static corpus a rival recreates with synthetic data and a year of usage?
When the model writes most of the code, what's the highest-value thing a human engineer still does?
Reliability guarantees in production AI
Reasoning proposes, infrastructure enforces. Everyone agrees on the principle. Nobody agrees where the boundary sits.
Doors at 1PM. Last call at 9PM.
1:00 – 2:00 PM
Doors open · Welcome
Check-in, coffee, and informal networking.
2:00 – 2:15 PM
Opening framing
Nick Rudder, CEO, Sphere: the infrastructure thesis behind AI in heavily regulated industries.
2:15 – 3:45 PM
Panel 1
Where do reliability guarantees sit in compliance heavy industries and where does automation stop and human review begin?
3:45 – 4:15 PM
Subsurface session
A 20-minute deep dive into a problem submitted with an application, tied to Panel 1.
4:15 – 5:15 PM
Break
Introductions matched across the room. Coffee, light bites.
5:15 – 6:30 PM
Panel 2
Is deep domain depth genuine technical defensibility, or does it just delay the inevitable as models climb?
6:30 – 7:00 PM
Subsurface session
The second deep dive from the room.
7:00 – 9:00 PM
Reception
Heavy appetizers, open bar, and networking until late.
Location
Dogpatch,
San Francisco
Doors at 1PM. Last call at 9PM. Check-in, coffee, and informal networking.
Dogpatch, San Francisco
The practical part
Yes. Ground Truth is hosted by Sphere. Your application is the ticket.
Coffee and light bites through the afternoon, heavy appetizers and an open bar at the reception. Come hungry, stay late.
Senior engineers and technical leaders building AI where regulation bites: tax, finance, mortgages, payroll, insurance, and their neighbors. If a wrong answer in your product means a legal or financial consequence, this room is for you.
Invites are individual, but have them apply. Teammates working on the same systems tend to make strong applications.
No. Panels and Subsurface sessions stay in the room.