Protocol

Triton Inference Protocol

KServe · V2 inference protocol

Lab ready Full lab

HTTP/gRPC model-inference APIs popularized by Triton/KServe V2: readiness/liveness, infer, and model-repository control. Tensor contents, datatypes, and shared-memory extensions dominate latency. It standardizes serving regardless of the backend framework.

Domains

Tags

inferenceservingtensorsml

Sources

  • KServe V2 inference protocol
  • NVIDIA Triton documentation

Related protocols