all docs

ASIC kernel offload

Ports deterministic inference kernels to specialized non-CUDA LPUs.

problem it solves
Avoids high CUDA GPU cloud rental costs for fixed deterministic subroutines.

What it does

What it does: Compiles and offloads fixed inference subroutines to non-CUDA ASIC hardware.

What it watches: Model kernel graphs, deterministic execution stages, and ASIC provider availability.

When it triggers: On requests targeting compatible deterministic model operations.

The Action: Dispatches matrix operations to specialized Groq LPUs or Cerebras hardware.

How it recovers: Standard CUDA GPU endpoints serve traffic if ASIC capacity is unavailable.

What we need from you

  • Attached Groq or Cerebras accountrequired

    Requires API credentials or hardware connection configured in control plane.

  • Compatible model architecturerequired

    Target model subroutines must compile to LPU/ASIC target.

What each mode does

ModeEffect on your requestWhat you can see
offNot consulted. All workloads execute on standard CUDA GPUs.No asic_offload stage recorded.
shadowVerifies ASIC kernel compilation and measures hypothetical cost delta.Stage with action=would_offload_asic logged.
prodRoutes compiled subroutines to ASIC hardware endpoints.ASIC tok/s throughput and cost savings logged.

Current policy

Supported targetsGroq LPU, Cerebras CS-3High-throughput deterministic hardware.
Cost efficiencyUp to 5x tokens per dollarCompared to standard cloud CUDA GPU rental.
Latency profileUltra-low TTFT (<50ms)Deterministic SPU/LPU execution.

Worth knowing before you enable it

  • ·Non-standard custom CUDA C++ kernels require porting to ASIC compiler toolchains.
  • ·ASIC architectures excel at linear deterministic sequence execution.
  • ·Requires provider credentials attached to ACE control plane.

What it replaces

  • ·High-cost CUDA GPU cloud rentals for deterministic workloads.
  • ·Manual model rewrites for specialized hardware.
  • ·Vendor-specific SDK lock-in code.

Dedicated ASIC hardware allocation and custom compiler passes on enterprise.

  • ·Custom ASIC compiler optimization passes.
  • ·Dedicated LPU hardware rack integration.
  • ·Multi-ASIC fallback topology routing.
team@acefleet.dev →