dotcore
Platform

On-premises AI inference, operated as a service.

NVIDIA GPU nodes, a managed software stack, and remote operations. Deployed as a portable unit or a containerized cluster on your network. We run it. Your data and models stay on hardware you own.

01 · HARDWARE

Two form factors, one software stack.

Build on the portable unit and scale to a cluster without changing your application.

Portable
  • 2 to 4 NVIDIA Blackwell-generation GPUs
  • Up to 384 GB VRAM; 70B models on a single GPU
  • ~2.9 to 5.8 PFLOPS FP8
  • Standard 15 to 20A wall outlet, ~1 kW
  • Operational 60 seconds after power-on
Container · 8 ft enclosure
  • 2 to 6 nodes, 16 to 48 GPUs
  • Up to ~4.6 TB VRAM; ~23 to 70 PFLOPS FP8
  • Power, cooling, fire suppression, physical security integrated
  • 100A 208V 3-phase; commissioned in 12 to 14 weeks
  • Larger containerized builds beyond 100 PFLOPS available
02 · SOFTWARE

The same managed stack on every unit.

Built on standard, open components. No proprietary runtime to lock you in. A full inference platform, not just a model server, from the API your apps call down to the silicon.

Applications & APIs
Your services call standard HTTP and gRPC endpoints.
Orchestration & routing
Requests routed across models and nodes from one entry point.
Guardrails & policy
Input and output filtering, red-teamed before rollout.
Inference serving
Models compiled for the GPUs they run on.
Models
Open weights, right-sized per workload.
Data & memory
Vector store, cache, and fault-tolerant NVMe across nodes.
Identity & secrets
Access control and key management you administer.
Operations & lifecycle
Telemetry, updates, and a control-plane agent we run remotely.
Hardware
NVIDIA GPU nodes, on your network, behind your firewall.

The specifics on every unit:

03 · MODELS

Open weights, on hardware you own.

04 · OPERATIONS & DATA

We run it. Nothing leaves your building.

05 · DELIVERY

From proof of concept to production.

A demo is easy; an AI system your business runs on is hard. We build every step against the production target, so nothing gets thrown away, and the whole path takes months, not years.

01 · Proof of concept

Prove it on real data.

We stand up your highest-value use case on a portable unit and validate it against your own private data, with no cloud round-trip.

Days to first inference
02 · Pilot

Harden and measure.

Right-size the models, add guardrails, wire it into your stack, and measure cost and latency under real load.

Weeks
03 · Production

Scale on owned hardware.

The same stack scales from a single unit to a containerized cluster. We operate it remotely; you own it outright.

In production within ~6 months

On-prem inference, operated as a service.

Talk to us