all docs

Low-precision quantization

FP8 and AWQ 4-bit weight quantization engine for GPU inference.

problem it solves
Prevents VRAM bottlenecks from limiting serving batch sizes and throughput.

What it does

What it does: Converts 16-bit floating point model weights to FP8 or 4-bit AWQ representations.

What it watches: Model weight precision, VRAM footprint, and numerical accuracy benchmarks.

When it triggers: When loading models configured for low-precision inference execution.

The Action: Halves or quarters VRAM consumption, doubling maximum serving batch capacity.

How it recovers: Accuracy validation suites verify model output against unquantized baselines.

What we need from you

  • Self-hosted model with configurable weightsrequired

    Requires control over model deployment artifacts and engine.

  • TensorRT-LLM or vLLM AWQ/FP8 supportrequired

    Hardware must support target low-precision instruction sets.

What each mode does

ModeEffect on your requestWhat you can see
offNot consulted. Models run in standard FP16 or BF16 precision.No quantization stage recorded.
shadowEvaluates quantized model outputs in background against FP16 baseline.Stage with action=would_quantize and perplexity_delta logged.
prodServes inference using low-precision quantized model build.VRAM footprint and batch size capacity logged.

Current policy

Supported formatsFP8 (E4M3/E5M2), AWQ 4-bit, INT8Hardware-accelerated quantization.
VRAM saving50% (FP8) to 75% (INT4)Frees memory for larger batch sizes.
Throughput gain1.5x - 2.2x batch throughputReduces cost per generated token.

Worth knowing before you enable it

  • ·FP8 requires modern GPU architecture (NVIDIA Ada Lovelace / Hopper / Blackwell).
  • ·Extremely small 4-bit models require activation-aware quantization (AWQ) to preserve reasoning.
  • ·Quantized weight builds must be pre-compiled or loaded from registry.

What it replaces

  • ·Manual post-training quantization conversion scripts.
  • ·High VRAM hardware requirement bills.
  • ·Single-request batch size memory limitations.

Custom calibration datasets and automated quantization build pipelines on enterprise.

  • ·Custom activation calibration dataset generation.
  • ·Automated TRT-LLM engine building.
  • ·Per-layer mixed-precision quantization schemes.
team@acefleet.dev →