Research lab · Self-deployed frontier models

Frontier models. Your hardware.

Make self-deploying open-source models cheaper, easier and faster. We choose, adapt and deploy the model around your workload and the cards you can afford.

For product teams using paid model APIs in production, operators running agents around the clock, and companies whose data must stay on premises. For teams already running open models or preparing to deploy them on their own hardware or cloud account.

What are frontier models? A term for models considered among the most capable. Their fit for your task still needs testing. A guide to the terms used here.

Choose the work you need

Start with the decision or build in front of you.

See all services

Research you can inspect

The lab researches engines, quantization, pruning, speculative decoding and Hebrew models. Published experiments describe their own setup and limits; they are not promises about your deployment.

Browse the research

Shared scale from zero; units: %

Recorded code sequences

Recorded code sequences: BF16 62.61 %; NVFP4 61.32 %BF1662.61 %NVFP461.32 %
First-token draft agreement on recorded code sequences. (%)Measured on our reference setup on Hardware: RTX PRO 6000 BlackwellQwen3.5-9B on fixed code sequences: the percentage of first proposed tokens matching the main model. A token is a text unit, such as part of a word. BF16 is the original weight format tested; the NVFP4 version stores some weights with fewer bits. This measures neither answer quality nor text generation speed.Receipt hqmtpUnderstand the terms and metrics in this chart
Inspect the data table
First-token draft agreement on recorded code sequences. · %
ConditionBF16NVFP4
Recorded code sequences62.61 %61.32 %

Quantization reduces the memory needed to store the model’s numbers and can make it possible to run on a card with less memory. This chart tests an acceleration method that proposes a token, a unit of text, before the main model verifies it. First-token agreement was similar with and without quantization in this experiment. This does not measure answer quality. Learn about tokens and speculative decoding.

What fits your hardware?

A model fitting in memory is only the start. The choice also depends on answer quality, request size, simultaneous work and the cost of operating the system.

Task and examplesQuality checksExpected loadHardware and budgetDecision factorsModel choiceMemory and executionResponse timeTotal operating costOutputA measuredrecommendation and adeployment scope
  1. Inputs

    • Task and examples
    • Quality checks
    • Expected load
    • Hardware and budget
  2. Decision factors

    • Model choice
    • Memory and execution
    • Response time
    • Total operating cost
  3. Output

    • A measured recommendation and a deployment scope.

Discuss an assessment

Does the move pay?

We compare your API bill with the full cost of running a suitable open model on your hardware or cloud account. The proposal covers adaptation, deployment, infrastructure, operation and support, with setup and running costs shown separately. It shows whether the move saves money. You control the deployment, data and logs.

From examples to a system you can run

Each stage has a clear purpose, from the first comparison to handover and continuing support.

  1. Choose

    Define the task and compare model and hardware options.

  2. Adapt

    Test fine-tuning, architecture and engine changes where they address a measured gap.

  3. Deploy

    Install, test and document the system on your hardware or cloud account.

  4. Support

    Support coverage and handover are set out in your proposal.

Our engineers keep supporting the systems they build, under agreed terms. About the lab →

Bring the workload and the budget.

We will work out what to measure and what to build.