Senior Software Engineer

Nikilesh K

Senior Software Engineer specializing in custom operator development, kernel optimization, and LLM/VLM bring-up on targeted hardware accelerators.

I bring hardware-aware ML systems engineering - from kernel-level optimization to full model inference deployment on custom targets.

4+
Years Experience
6+
Frameworks Mastered
3
DSP Platforms
5+
Inference Servers

Core Competencies

Deep expertise spanning the full ML systems stack - from hardware-level kernel tuning to production inference serving.

C / C++ Python Bash / Shell PyTorch Transformers vLLM SGLang ONNX / ONNX-RT NumPy / OpenCV Custom Op Development Kernel Design & Opt. SIMD / VECC Kernels DMA & Memory Opt. DSP Acceleration LLM / VLM Bring-Up Model Inference Model Conversion AI Agentic Workflows Kubernetes / GCP Docker / CI-CD Profiling & Debugging MLA Attention Decoupled-Rope MTP

What I Build

Evidence-first engineering - measured impact, not just claims.

Radar-SDK Porting & Development DSP / Custom

Ported and optimized Radar modules on Custom DSP Platform - FFT, CA-CFAR, OS-CFAR, DML single/dual target detection with full SDK framework integration.

3
DSP Platforms
5+
Radar Modules
100%
SDK Integration
SIMD & VECC Kernel Optimization Low-Level

Developed optimized SIMD/VECC kernels for Custom DSP workloads to maximize compute efficiency and hardware utilization - targeting cycle-level performance gains.

Max
Hardware Util.
Cycle
Level Tuning
DMA & Memory Optimization Performance

Implemented DMA-based data transfers with single/double buffering to reduce latency and improve memory efficiency - analyzed theoretical vs. achieved cycles for vector kernel optimization.

2×
Buffer Strategy
↓
Latency Reduced
Custom Device Operator & Kernel Development C++ / Python

Built custom device operators for targeted hardware - native convolution, deformable convolution, and various categories of kernels such as arithmetic, elementwise, comparison, logical kernels optimized for latency, throughput, and memory.

3
Kernel Types
Native
Execution
ONNX Custom Operator Development ONNX-RT

Designed and integrated custom ONNX operators into ONNX Runtime (x86) - optimized convolution kernels, registered custom ops, validated correctness and performance. Built a symbolic-mapping linker bridging PyTorch and ONNX.

PyTorch → ONNX
Bridge Built
✓
Validated
LLM / VLM Bring-Up & Inference Serving ML Systems

Built and optimized pipelines for large language models - multimodal and text-based - integrated into vLLM inference server for deployment with debugging and model-serving expertise. GCP-based cloud deployment with Kubernetes orchestration, containerization, and CI/CD pipelines for scalable inference serving.

vLLM
Inference
GCP
Cloud

Professional Journey

4+ years of engineering across ML systems, kernel optimization, and hardware acceleration.

Senior Software Engineer 4+ Years
Multicoreware Inc

Specialized in custom operator development, kernel design and optimization, LLM/VLM model bring-up on targeted hardware accelerators, runtime integration, and performance tuning. Deep knowledge of recent LLM architectures (MLA attention, Decoupled-Rope, MTP). Strong experience with heterogeneous compute platforms, DSPs, and custom accelerators. Hands-on with AI agentic workflows, vLLM/SGLang inference serving, and Google Cloud infrastructure (GKE, Cloud Storage). Experienced in Agile, CI/CD, containerization, and distributed team collaboration.

Academic Background

Master of Computer Applications (MCA) 2019 – 2022
Dwaraka Doss Goverdhan Doss Vaishnav College

Computer Programming, Specific Applications

Bachelor's Degree, Mathematics 2016 – 2019
Dwaraka Doss Goverdhan Doss Vaishnav College

Debugging Philosophy

  • Fusion vs. Scheduling Overhead - The real bottleneck in distributed inference isn't compute; it's kernel launch granularity and memory copy overhead. FFN stages are compute-bound; attention is memory-bandwidth-bound when decode.
  • Where Theoretical Speedups Are Won or Lost - DMA double-buffering hides transfer latency, but only when pipeline depth matches the accelerator's command queue depth. Profile, don't guess.
  • Cross-Framework Bridge - Symbolic mapping between PyTorch and ONNX is where numerical precision silently degrades. Validate at every boundary.
  • Model Bring-Up - On-device LLM deployment is a memory-constrained problem first, a compute problem second. Quantization-aware design from day one.

Full-Stack ML Engineering

Model Optimization · Device Ops & Kernel Optimization · Model Enablement · Kernel Development · Inference Platform Integration · AI Agentic Automation · Hardware Acceleration (DSPs, Custom Accelerators) · Distributed ML Systems · Cloud Infrastructure (GCP/GKE) · CI/CD & Containerization

Let's Build Something

Always interested in challenging ML systems work and hardware-accelerated inference.