Low-precision quantization
FP8 and AWQ 4-bit weight quantization engine for GPU inference.
problem it solves
Prevents VRAM bottlenecks from limiting serving batch sizes and throughput.
What it does
What it does: Converts 16-bit floating point model weights to FP8 or 4-bit AWQ representations.
What it watches: Model weight precision, VRAM footprint, and numerical accuracy benchmarks.
When it triggers: When loading models configured for low-precision inference execution.
The Action: Halves or quarters VRAM consumption, doubling maximum serving batch capacity.
How it recovers: Accuracy validation suites verify model output against unquantized baselines.
What we need from you
- Self-hosted model with configurable weightsrequired
Requires control over model deployment artifacts and engine.
- TensorRT-LLM or vLLM AWQ/FP8 supportrequired
Hardware must support target low-precision instruction sets.
What each mode does
| Mode | Effect on your request | What you can see |
|---|---|---|
| off | Not consulted. Models run in standard FP16 or BF16 precision. | No quantization stage recorded. |
| shadow | Evaluates quantized model outputs in background against FP16 baseline. | Stage with action=would_quantize and perplexity_delta logged. |
| prod | Serves inference using low-precision quantized model build. | VRAM footprint and batch size capacity logged. |
Current policy
| Supported formats | FP8 (E4M3/E5M2), AWQ 4-bit, INT8 | Hardware-accelerated quantization. |
| VRAM saving | 50% (FP8) to 75% (INT4) | Frees memory for larger batch sizes. |
| Throughput gain | 1.5x - 2.2x batch throughput | Reduces cost per generated token. |
Worth knowing before you enable it
- ·FP8 requires modern GPU architecture (NVIDIA Ada Lovelace / Hopper / Blackwell).
- ·Extremely small 4-bit models require activation-aware quantization (AWQ) to preserve reasoning.
- ·Quantized weight builds must be pre-compiled or loaded from registry.
What it replaces
- ·Manual post-training quantization conversion scripts.
- ·High VRAM hardware requirement bills.
- ·Single-request batch size memory limitations.
Custom calibration datasets and automated quantization build pipelines on enterprise.
- ·Custom activation calibration dataset generation.
- ·Automated TRT-LLM engine building.
- ·Per-layer mixed-precision quantization schemes.