Nikilesh K
Senior Software Engineer specializing in custom operator development, kernel optimization, and LLM/VLM bring-up on targeted hardware accelerators.
I bring hardware-aware ML systems engineering - from kernel-level optimization to full model inference deployment on custom targets.
Core Competencies
Deep expertise spanning the full ML systems stack - from hardware-level kernel tuning to production inference serving.
What I Build
Evidence-first engineering - measured impact, not just claims.
Ported and optimized Radar modules on Custom DSP Platform - FFT, CA-CFAR, OS-CFAR, DML single/dual target detection with full SDK framework integration.
Developed optimized SIMD/VECC kernels for Custom DSP workloads to maximize compute efficiency and hardware utilization - targeting cycle-level performance gains.
Implemented DMA-based data transfers with single/double buffering to reduce latency and improve memory efficiency - analyzed theoretical vs. achieved cycles for vector kernel optimization.
Built custom device operators for targeted hardware - native convolution, deformable convolution, and various categories of kernels such as arithmetic, elementwise, comparison, logical kernels optimized for latency, throughput, and memory.
Designed and integrated custom ONNX operators into ONNX Runtime (x86) - optimized convolution kernels, registered custom ops, validated correctness and performance. Built a symbolic-mapping linker bridging PyTorch and ONNX.
Built and optimized pipelines for large language models - multimodal and text-based - integrated into vLLM inference server for deployment with debugging and model-serving expertise. GCP-based cloud deployment with Kubernetes orchestration, containerization, and CI/CD pipelines for scalable inference serving.
Professional Journey
4+ years of engineering across ML systems, kernel optimization, and hardware acceleration.
Specialized in custom operator development, kernel design and optimization, LLM/VLM model bring-up on targeted hardware accelerators, runtime integration, and performance tuning. Deep knowledge of recent LLM architectures (MLA attention, Decoupled-Rope, MTP). Strong experience with heterogeneous compute platforms, DSPs, and custom accelerators. Hands-on with AI agentic workflows, vLLM/SGLang inference serving, and Google Cloud infrastructure (GKE, Cloud Storage). Experienced in Agile, CI/CD, containerization, and distributed team collaboration.
Academic Background
Computer Programming, Specific Applications
Debugging Philosophy
- Fusion vs. Scheduling Overhead - The real bottleneck in distributed inference isn't compute; it's kernel launch granularity and memory copy overhead. FFN stages are compute-bound; attention is memory-bandwidth-bound when decode.
- Where Theoretical Speedups Are Won or Lost - DMA double-buffering hides transfer latency, but only when pipeline depth matches the accelerator's command queue depth. Profile, don't guess.
- Cross-Framework Bridge - Symbolic mapping between PyTorch and ONNX is where numerical precision silently degrades. Validate at every boundary.
- Model Bring-Up - On-device LLM deployment is a memory-constrained problem first, a compute problem second. Quantization-aware design from day one.
Full-Stack ML Engineering
Model Optimization · Device Ops & Kernel Optimization · Model Enablement · Kernel Development · Inference Platform Integration · AI Agentic Automation · Hardware Acceleration (DSPs, Custom Accelerators) · Distributed ML Systems · Cloud Infrastructure (GCP/GKE) · CI/CD & Containerization