Remote: Yes
Willing to relocate: Yes (India, Europe, US)
Technologies: Python, C++, PyTorch, FastAPI, Docker, Playwright, LLM agents (ReAct, MCP, tool calling), RAG, QLoRA/PEFT, quantization, Redis, PostgreSQL, AWS
Résumé/CV: https://drive.google.com/file/d/1qUxHkYu8dx866aGJy4-GLejJui7...
Email: shrey7shrey@gmail.com
Recent CS grad. I build agent systems and the verification machinery that proves they actually worked — the second part is where most of my time goes.
KARMA (github.com/ShreySharma07/Multi-Agentic-CUDA-optimization): 5-agent pipeline where LLM agents read NVIDIA Nsight Compute traces and rewrite CUDA kernels from measured bottlenecks. 8.99x over PyTorch eager, 87.9% of theoretical peak memory bandwidth, on a 6GB consumer GPU. I got a 40x early on that turned out to be a kernel computing garbage very fast — the model was optimizing my metric, not the task. Most of the work after that was a verification harness it couldn't game: compile gate, 3-seed correctness vs PyTorch, determinism gate, interleaved benchmarking. 0% false positives across every number I report.
Repliq (github.com/ShreySharma07/salesforce-agent): autonomous computer/browser-use agent. One narrated screen recording in, self-verifying execution plan out then the agent performs the task in a secured docker environment. 20-step Salesforce workflow end to end. Dominant failure mode was steps reporting success while the record silently never saved — fixed with per-primitive postconditions checked against real application state instead of agent self-report. just connect it with your org give me the task recording and done you don't have to worry about the day to day redundant tasks.
Looking for AI/ML or backend engineering. Happy to go deep on either project.