Manages memory-efficient model inference by automatically batching inputs
inference_manager.RdService class that orchestrates memory-efficient inference through automatic batching, flexible offloading (GPU/CPU/disk), and OOM recovery.
inference_manager.RdService class that orchestrates memory-efficient inference through automatic batching, flexible offloading (GPU/CPU/disk), and OOM recovery.