Tech Week Singapore 2026
Rethinking the Data Path for AI Inference
29 Sept 2026
Cloud & AI Infrastructure Keynote Theatre
As generative AI moves from experimentation into production, inference is becoming the defining challenge for enterprise AI. Adding more GPUs alone cannot solve rising latency, concurrency, and token-cost pressures; the way data and KV cache move across the infrastructure stack is increasingly critical. In this session, Nikhil Madan will examine how a tiered approach spanning GPU memory, CPU memory, local NVMe, and shared high-performance storage can remove inference bottlenecks. Attendees will learn how better KV-cache orchestration can reduce time to first token, increase throughput and concurrency, improve GPU utilization, and lower the cost of serving AI at without continually expanding costly GPU infrastructure.
Cloud & AI Infrastructure
Cyber Security World
Big Data & AI World
Data Centre World 
























