AI Infrastructure & Software Engineer

Vineet Agarwal

AI Infrastructure · Backend · Full-Stack Systems

I build AI systems and the infrastructure that runs them. My recent work spans LLM inference optimization, retrieval augmented generation, on-device AI, and multi agent workflows, backed by 2+ years shipping production web, mobile, and cloud native systems. I care about latency, throughput, and reliability, and I like owning products end to end, from frontend and backend APIs down to GPU level model serving and deployment.

Open to full-time software engineering roles starting January 2027

Portrait of Vineet Agarwal

Skills

Engineering across AI infrastructure, backend, mobile, and cloud

My experience spans model serving and applied AI through production APIs, cloud infrastructure, real-time systems, and cross-platform applications.

AI Infrastructure & Inference

Serving and optimizing LLMs for throughput, latency, and memory efficiency.

  • PyTorch
  • Hugging Face
  • vLLM
  • SGLang
  • TensorRT-LLM
  • KV Cache
  • Quantization
  • Speculative Decoding

Applied AI & Agents

RAG pipelines, retrieval, and multi-agent systems over real data.

  • LangChain
  • RAG Pipelines
  • FAISS
  • AutoGen
  • Google ADK
  • OpenAI APIs
  • Multi-Agent Systems
  • MCP Servers

Backend & APIs

Production APIs, real-time services, auth, and system design.

  • Python
  • FastAPI
  • Node.js
  • Express
  • REST APIs
  • WebSocket
  • Authentication
  • System Design

Cloud & Infrastructure

Containers, orchestration, CI/CD, and cloud-native deployment.

  • AWS
  • Docker
  • Kubernetes
  • GitHub Actions
  • CI/CD
  • CloudFormation
  • Monitoring
  • Linux

Frontend & Mobile

Cross-platform product interfaces on web, iOS, and Android.

  • React
  • React Native
  • Next.js
  • TypeScript
  • Redux Toolkit
  • Tailwind CSS
  • Android
  • iOS

Data & Performance

Relational and NoSQL data, caching, and query optimization.

  • PostgreSQL
  • MongoDB
  • MySQL
  • Redis
  • SQL
  • Data Modeling
  • Query Optimization
  • Caching

Background

Professional experience and education

Selected Work

Projects I have built and shipped

A mix of on-device AI, full-stack products, and real-time systems — spanning inference, backend, cloud, and mobile.

Locra : Private AI Without Wi-Fi project preview

Locra : Private AI Without Wi-Fi

100% on-device inference

Privacy-first React Native assistant that performs text and image Q&A entirely on-device using Qwen3-VL through llama.rn. It includes local model downloading and verification, streamed responses, persistent conversations, cancellation, and memory-aware failure handling without sending user content to a server.

  • React Native
  • TypeScript
  • llama.rn
  • Qwen3-VL
  • On-device AI
CareBridge AI project preview

CareBridge AI

Earlier risk detection

AI powered care transition platform that turns discharge PDFs into structured clinical data and actionable care plans. Built with FastAPI, React, PostgreSQL, and Google ADK, using multi agent validation and confidence threshold guardrails to catch post discharge risks earlier.

  • FastAPI
  • Google ADK
  • AI Agents
  • React
AI Repository Analyzer project preview

AI Repository Analyzer

3 specialized agents

Multi agent code analysis platform that ingests GitHub repositories or ZIP files and returns architecture insights, API understanding, and context aware Q&A. Built with FastAPI, LangChain, AutoGen, FAISS, and OpenAI, using a RAG layer over the codebase and specialized SDE, PM, and QA agents.

  • AutoGen
  • LangChain
  • RAG
  • FAISS
AWS Video Analytics Streaming Platform project preview

AWS Video Analytics Streaming Platform

0 critical vulnerabilities

Team project · Backend and cloud contributor

Cloud native, serverless analytics platform on AWS using Lambda, EKS, S3, DynamoDB, SQS, API Gateway, and CloudFormation. Microservices auto scale from 2 to 10 pods behind a load balancer, with infrastructure as code and a GitHub Actions CI/CD pipeline that passed security scanning with zero critical vulnerabilities.

  • AWS Lambda
  • EKS
  • Serverless
  • CloudFormation
Terrapin Events - Campus Event Management project preview

Terrapin Events - Campus Event Management

Sub 200ms at 100+ users

Team project · Full-stack and infrastructure contributor

Full stack campus event platform with CAS SSO, JWT role based access control, Stripe payments, and waitlist automation. Built with React, FastAPI, and MongoDB and deployed on Kubernetes with CI/CD, holding sub 200ms API performance under 100+ concurrent users.

  • React
  • FastAPI
  • Kubernetes
  • MongoDB
Text to Image Llama project preview

Text to Image Llama

100% local inference

Local AI image generation system that strengthens weak prompts through an LLM powered enhancement layer before diffusion inference. Built with FastAPI, llama.cpp, and diffusers to combine local model serving, prompt refinement, and modular backend orchestration without relying on hosted APIs.

  • FastAPI
  • llama.cpp
  • Local Serving
  • Diffusers

Contact

Open to AI infrastructure, backend, and full-stack software engineering roles

If you are hiring or want to talk through a problem in AI inference, backend systems, or full-stack product work, I would love to hear from you.

Let's connect

Whether it's a full-time role, a technical discussion, or a collaboration, feel free to reach out.

Best fit for this portfolio

AI infrastructure and LLM inference, applied AI and agents, backend and APIs, cloud-native systems, and full-stack and mobile product work.

Send a message