The Code Jeweller ← The atelier
I · The Floor

Every evaluation your system runs is a cost. I bring that cost down.

For principals whose systems already think for a living. If you came for a digital house, the Work below is enough — this section is the metal underneath.

A vetting system reads, weighs, and scores every opportunity it's shown — and most operators pay retail to do that, on someone else's hardware, at someone else's margin. Lower the cost per evaluation and the same budget covers more opportunities: a wider net, a harder filter, a sharper signal out the other end. This is the same discipline as the rest of the practice, applied one layer deeper.

Method

Compression, routing, self-hosting

Instruction compression, context shaped to be read in fewer tokens, live routing across models most operators never track, free-tier capacity mapped and used, and inference run on owned hardware where it makes sense. Named here; the specifics stay mine — that's the edge you'd be retaining.

Named as craft, not explained as method
Economics

Fixed cost, not metered

Self-hosted inference means the marginal cost of one more evaluation trends toward electricity, not a per-token bill that grows with your ambition. Scale stops being a tax.

Their cost is variable. Mine is fixed.
Delivery

One point of judgment, start to finish

The same person who engineers the inference builds what runs on top of it. No handoff, no committee, no translation loss between the two.

What's designed is what ships
Sovereignty

What you build stays exactly where you put it

Nothing rented, nothing borrowed, nothing another company's policy can quietly change underneath you. Your intelligence lives on hardware only you can reach. A separate, deeper practice for principals who need it — ask, if that's the concern.

Savings, and sovereignty

Give me one real slice of your vetting workflow. I'll run it two ways — your current cost, and mine — on your own numbers, and you decide from there.

II · The Bench

The workshop, visible.

Every claim on this page runs on hardware I own. Not a stock photo of a server room — the actual bench: three nodes, three architectures, one pull-queue. A snapshot of real measurements from the working mesh, dated like anything honest.

Measured · 60-item job

6.0 s across the mesh. 19.1 s on the best single node.

Six local workers across three machines finished the same job 3.2× faster than the strongest node alone — accuracy equal or better, scored against mechanical ground truth. The finding that matters: no routing logic. Workers pull from one queue; the fast node simply takes more work. Don't build a router. Build a queue.

Self-balancing by construction
Resident · verify tier

A 27B-class model, 32k context, fully on owned silicon

Drafts run small and parallel; a larger local model reviews them; frontier intelligence is spent only where judgment is genuinely needed. The same economics I sell — running underneath this very site's production.

The method, eating its own cooking
Discipline

Local first, when it doesn't cost quality

Bulk work — classification, drafting, conversion, screening — runs on the bench at the marginal cost of electricity. Latest cut, measured on a 326,531-file classification run: 91% of each item's compute was the model re-reading identical prompt boilerplate. Trimmed it, re-scored it, accuracy unchanged within noise — the whole mesh got faster by deleting words, not buying hardware. Knowing what to delete is the craft.

Snapshot · 21 July 2026
Seen enough metal?
Request an assay
No CMS, no framework, no cookies — hand-cut HTML and three machines arguing quietly over one queue. The cheap work costs electricity. The expensive work is judgment. Every number wears a date.
Ora et Labora · Marbella · Gerard Aparicio Oliart