Soham Dutta
ML systems engineer working on the infrastructure underneath large models — inference serving, distributed training, and memory management.
Signal processing background (IIEST Howrah) moving down the stack into systems: KV cache management, communication-compute overlap, and inference engine internals.
Background
I started in signal processing and stochastic modeling — MDPs, SDEs, and Monte Carlo methods for financial risk at IIT Bombay's FedEx Centre for Advanced Logistics, and quantum simulation research at the Indian Statistical Institute. That work was fundamentally about modeling systems under uncertainty at scale.
I've since moved toward the systems layer of ML: how inference actually gets served, how gradients get communicated across GPUs, and how memory gets managed under load. My current projects sit at the intersection of compilers, distributed systems, and ML serving infrastructure.
Present
Jul 2025
Experience
- Developed MDPs and MCMC methods for financial risk analysis, improving simulation accuracy by 20%.
- Ran 200+ simulations analyzing random walk convergence to GBM via Donsker's Invariance Principle, increasing prediction reliability by 25%.
- Modeled asset prices with SDEs achieving 95% accuracy.
- Implemented Monte Carlo simulated annealing to optimize MUBs in 4D quantum systems, reducing error rate 20% vs. traditional methods.
- Simulated up to 4 MUBs in 4D quantum systems achieving 98% measurement fidelity.
- Optimized production output for an industrial refinery client, reducing operational costs by 10%.
- Built regression models using SGD & genetic algorithms, maintaining error margin under 3%.
Systems & infra projects
Active portfolio work, ordered by what I'm deepest into right now.
A vLLM PagedAttention-inspired memory manager built from scratch: block allocator, block table, and copy-on-write semantics for beam search, split across a three-crate Rust workspace with a CUDA kernel layer.
Custom PyTorch autograd hooks with NCCL all-reduce bucketing to overlap gradient communication with backward compute in distributed training. Runs locally over gloo, benchmarked on rented multi-GPU cloud instances.
Benchmarks RadixAttention against multi-turn agent sessions, multi-agent fan-out, RAG pipelines, and self-consistency sampling on a single rented GPU, instrumented with Prometheus/Grafana, tracking cost-per-1,000-agent-turns.
A three-tier memory architecture (hot / warm / cold) with salience scoring and KV cache eviction, served through a vLLM + LMCache layer — built to move past typical vector-database chatbot patterns.
Multi-agent framework enabling autonomous agents to exchange structured messages, negotiate tasks, and collaboratively solve problems via an asynchronous message-passing protocol. Part of IIEST Shibpur ETCE departmental projects.
Matmul kernel optimization using the MLIR Transform dialect, documented with a paper-style README.
A Graph Variational Autoencoder for generating molecular structures from graph-based chemical compound representations.
Stack
Systems & infra
Languages
ML frameworks
Platforms & tools
Get in touch
Open to ML infra, inference optimization, and systems engineering roles. Happy to talk about any of the projects above in detail.