Labs · prototype
We fund our own research,
then productise it
Agentic prototypes, permission models and evaluation methods. What survives Labs becomes client-grade.
Questions a client can't pay for
Labs exists because the interesting questions in agentic systems are not answerable on a client's budget. How much authority can an agent hold before a user stops trusting it? What does a reversible action actually require underneath? How do you evaluate a multi-step plan when only the final state is observable?
We fund that work ourselves, publish what we learn, and move the techniques that hold up into client engagements. Nothing reaches a client system until it has survived here first.
From prototype to every engagement
Scoped tool sandboxes. The least-privilege tool model Ella runs on is the standard for every client agent we build.
Evaluation from real cases. The method for turning a client's history into an eval set that decides whether the system ships.
Deterministic replay. Traces that reproduce a run bit-for-bit — first built to debug Ella, now the standard for every pipeline we deliver.
Ella acts for you — only where you allow it
A permission-based agent that carries out real tasks on your behalf. It asks before it acts, works inside the scopes you grant, and keeps a reversible ledger of everything it did. Applied agentic AI — custom orchestration and permissioned on-device actions, not a black box.
Ella, today
Where the prototype actually stands. We would rather say “prototype” and mean it than say “product” and not.
- Stage
- Prototype
- Availability
- Internal — not publicly available
- Runs on
- Local machine; actions execute on-device
- Scopes
- Per-task, plain-language, granted before any action runs
- Ledger
- Sealed, ordered, replayable; reversible for 30 days
- As of
- September 2026
public/images/ and pass src. The duotone treatment applies automatically.public/images/ and pass src. The duotone treatment applies automatically.Five things it will not skip
Each of these started as a Labs question and became a rule. Together they are why Ella can be trusted with real tasks.
Scopes, granted up front
Ella cannot touch anything you have not named. Scopes are requested per task, shown to you in plain language, and granted or refused before a single action runs. There is no 'allow everything' switch.
Ask before acting
Any step that changes state pauses for confirmation unless you have pre-approved that class of action for that scope. The default is to interrupt, not to guess.
Every action in a ledger
Each read, write and tool call is written to a sealed, ordered ledger with the inputs it saw. You can open any run, weeks later, and see exactly what happened and why.
Reversible for thirty days
Actions are journaled with enough state to undo them. For thirty days, anything Ella did can be rolled back — one action, or the whole session.
On-device where it matters
Actions that touch your files, your accounts or your data execute locally. What leaves the machine is the plan, not the payload.
A system, brought up
Every engagement ends with something running: typed steps, scoped tools, a confidence floor that escalates rather than guesses, and a ledger you can replay.
system online · every action traced
Let's build
the impossible.
Tell us the system you can't get built. We come back with a short, paid discovery — a clear plan and a fixed first milestone — usually within two working days.
Discovery
A working session to map the problem and define what "good" is measured against.
The plan
Architecture, milestones and a fixed first deliverable — yours to keep, either way.
We build
Embedded with your team or as a dedicated pod, shipping with traces, evals and docs.