0x00 — boot

Soham Dutta

ML systems engineer working on the infrastructure underneath large models — inference serving, distributed training, and memory management.

Signal processing background (IIEST Howrah) moving down the stack into systems: KV cache management, communication-compute overlap, and inference engine internals.

M.Tech, Signal Processing — IIEST Shibpur Aug 2025 – Present
Summer Research Intern — IIT Bombay, FedEx CALA Apr – Sep 2024
Research Associate — Indian Statistical Institute Mar 2024 – Apr 2025
Data Science Intern — Entiovi Technologies Jun – Sep 2023
B.Tech, Computer Science — Heritage Institute of Technology Oct 2021 – Jul 2025
0x01 — about

Background

I started in signal processing and stochastic modeling — MDPs, SDEs, and Monte Carlo methods for financial risk at IIT Bombay's FedEx Centre for Advanced Logistics, and quantum simulation research at the Indian Statistical Institute. That work was fundamentally about modeling systems under uncertainty at scale.

I've since moved toward the systems layer of ML: how inference actually gets served, how gradients get communicated across GPUs, and how memory gets managed under load. My current projects sit at the intersection of compilers, distributed systems, and ML serving infrastructure.

Aug 2025 –
Present
M.Tech, Signal Processing — GPA 8.6/10
IIEST Howrah
Oct 2021 –
Jul 2025
B.Tech, Computer Science & Engineering — GPA 8.0/10
Heritage Institute of Technology, Kolkata
0x02 — experience

Experience

Summer Research Intern — IIT Bombay, FedEx Centre for Advanced Logistics & Analytics
Apr – Sep 2024 · Mumbai
  • Developed MDPs and MCMC methods for financial risk analysis, improving simulation accuracy by 20%.
  • Ran 200+ simulations analyzing random walk convergence to GBM via Donsker's Invariance Principle, increasing prediction reliability by 25%.
  • Modeled asset prices with SDEs achieving 95% accuracy.
Python · NumPy · SciPy · Pandas · Matplotlib
Research Associate — Indian Statistical Institute, Applied Statistics Unit
Mar 2024 – Apr 2025 · Kolkata
  • Implemented Monte Carlo simulated annealing to optimize MUBs in 4D quantum systems, reducing error rate 20% vs. traditional methods.
  • Simulated up to 4 MUBs in 4D quantum systems achieving 98% measurement fidelity.
Qiskit · PennyLane · D-Wave
Data Science Intern — Entiovi Technologies
Jun – Sep 2023 · Kolkata
  • Optimized production output for an industrial refinery client, reducing operational costs by 10%.
  • Built regression models using SGD & genetic algorithms, maintaining error margin under 3%.
Scikit-Learn · TensorFlow
0x03 — projects

Systems & infra projects

Active portfolio work, ordered by what I'm deepest into right now.

Paged KV Cache Memory Manager In progress

A vLLM PagedAttention-inspired memory manager built from scratch: block allocator, block table, and copy-on-write semantics for beam search, split across a three-crate Rust workspace with a CUDA kernel layer.

Rust · CUDA C++ · paged-kv-core · paged-kv-cuda
Communication–Compute Overlap Scheduler In progress

Custom PyTorch autograd hooks with NCCL all-reduce bucketing to overlap gradient communication with backward compute in distributed training. Runs locally over gloo, benchmarked on rented multi-GPU cloud instances.

PyTorch · NCCL · gloo · RunPod / Vast.ai / Lambda
SGLang / RadixAttention Benchmarking Suite In progress

Benchmarks RadixAttention against multi-turn agent sessions, multi-agent fan-out, RAG pipelines, and self-consistency sampling on a single rented GPU, instrumented with Prometheus/Grafana, tracking cost-per-1,000-agent-turns.

SGLang · Prometheus · Grafana
Stateful Multi-Session AI Agent

A three-tier memory architecture (hot / warm / cold) with salience scoring and KV cache eviction, served through a vLLM + LMCache layer — built to move past typical vector-database chatbot patterns.

vLLM · LMCache · Python
// earlier / academic
Agent-to-Agent Communication Framework

Multi-agent framework enabling autonomous agents to exchange structured messages, negotiate tasks, and collaboratively solve problems via an asynchronous message-passing protocol. Part of IIEST Shibpur ETCE departmental projects.

Python · LangChain · AutoGen · FastAPI · WebSockets · Redis
LLVM/MLIR Transform-Dialect Autotuning

Matmul kernel optimization using the MLIR Transform dialect, documented with a paper-style README.

LLVM · MLIR
Graph-VAE for Molecule Generation

A Graph Variational Autoencoder for generating molecular structures from graph-based chemical compound representations.

Python · PyTorch · TensorFlow
0x04 — stack

Stack

Systems & infra

CUDAvLLMSGLang NCCLRayDocker

Languages

PythonRustC++ JuliaSQLMATLABR

ML frameworks

PyTorchTensorFlowScikit-Learn QiskitPennyLane

Platforms & tools

AWSGCPLinux GitJupyter
0x05 — contact

Get in touch

Open to ML infra, inference optimization, and systems engineering roles. Happy to talk about any of the projects above in detail.