pebble · machine intelligence lab · san francisco

Machines learn by doing.

We build your self-improving superintelligence loop.

Self-improving intelligence cycle.

You name the outcome. We build the loop that reaches it: every stage feeds the next, and everything the loop produces stays on your side of the table.

OUTCOMES VERIFICATION ENVIRONMENT RL POST-TRAINING INFERENCE
01environment

An environment around your work.

We scope the outcome you need and build the environment that reaches it, however many processes that touches. It ships doing the job, on a frontier model to start.

Start a build →
02verification

Every step graded, on the record.

Checks you approve grade each run as it happens. Postmortems start from a step number, and reliability is a percentage read off the run history.

run 2141 · invoice-reconciliation · pebble-ir-8b
  step 03  match_po_number         pass
  step 04  amounts_reconcile       pass
  step 05  approval_chain_valid    pass
  step 06  ledger_entry_posted     pass
  outcome  approved · 11.2s · $0.003
trailing 30 days   2,141 runs · 98.7% approved
03training

An open model, trained on your approved runs.

The graded runs become training data. We post-train an open model on the runs your checks approved, on compute we schedule across clouds at the best rate that hour.

How compute works →
04inference

Cheaper per run, and yours.

The tuned model takes over the work at a fraction of frontier API price, behind the same checks. The records can retrain whatever base model comes next, so the cycle keeps compounding.

Talk to us →

Our clients' intelligence is their own.

What we can say about recent work:

audit · accounting

A large audit and accounting firm. An environment for AI-assisted audit work, with models post-trained and hosted on the full stack.

insurance

An AI insurance company. An environment built around their workflows, with post-training on the graded runs.

materials science

A large materials science company. Environment, post-training, and hosting across their research workflows.

drug · target discovery

Drug and target discovery companies. Environments built around their pipelines, with post-trained models served on the full stack.

cybersecurity

A cybersecurity company. Environment, post-training, and hosting, inside their perimeter.

fintech

A fintech company. An environment around their workflows, with post-training on the graded runs.

defense

A defense contractor. Models for drone intelligence at the edge, trained in an environment we built for it.

What an environment is.

An environment wraps software around your work. A machine does the steps on your systems, and approved verifiers grade each one.

Machines improve at whatever gets graded, so a machine in your environment improves at your process. The graded runs become training data, and a model trained on them has practiced your work toward the outcomes you want.

Reinforcement learning is an old idea. You improve at a trade by practicing it, and by knowing how each attempt turned out. Models improve the same way.

We build with you, around the capabilities you need and the outcomes you want, each one verifiable. Your machines improve on your work specifically, and the improvement is yours to keep.

Working with us.

2 · scope

You name the outcome. If an environment is the wrong tool for it, we say so.

1 · codesign

We design it with you and carry the build ourselves.

3 · run

Handoff, on your infrastructure or ours. We stay on for what comes next.

Never overpay for compute.

Compute is a commodity. Everyone buys from the same clouds and datacenters. The work is getting the right machines at the lowest price and running them well. That layer is ours, and we run it for you.

On demand

1 to 256 GPUs across clouds. One platform.

1.1
Scheduling

SLURM and Kubernetes, containers included.

1.2
Interconnects

Infiniband networking for training that spans nodes.

1.3
Observability

Grafana dashboards on every job, live.

Reserved clusters

Large clusters, quoted within a day.

2.1
Parallel bids

One request, priced across datacenters, H100 through GB300 NVL72.

2.2
Idle sell-back

Idle capacity resells on the spot market. Reclaim it the hour you need it.

2.3
An engineer on the line

The same person from cluster bring-up to steady state.

H100from $1.99/hr
80 GB VRAM · spot $0.94/hr
H200from $2.49/hr
141 GB VRAM · 44 vCPU
B200from $3.49/hr
192 GB VRAM · 32 vCPU
B300from $4.99/hr
288 GB VRAM · 48 vCPU
GH200from $3.14/hr
96 GB VRAM · 72 vCPU
A100from $1.19/hr
80 GB VRAM · 32 vCPU
one platform · metered per second · budget-capped

Recursively self-improving machine intelligence.

Rockie is an open platform where machine intelligence improves machine intelligence. Anyone with an idea and some conviction can run HPC science there and push the frontier.

We built it and maintain it for the public good, and it improves with every experiment that runs through it. An account takes a minute at rockielab.com.

We also run a research program.

Pebble publishes its own work on memory scaling laws and matrix state as memory. Code and eval data ship with every note, and negative results stay published.

Read the research log →

Contact.

[email protected]