Projects

What I've Built

01
01 / 05
2025 – Present
View Project

AIMO: Enterprise AI Agent Platform

No-code platform for building AI agents on any LLM, wired directly into the tools a workspace already runs on — Slack, Notion, GitHub, Salesforce, and 20+ other services — so agents act with full organizational context instead of a blank prompt.

20+
Integrations
Any
LLM Provider
Role-Based
Governance
Multi-LLMAgentsSlackNotionGitHubSalesforceRAGEnterprise
Read the case study
02
02 / 05
2024 – 2025

MedRAG: Clinical Dialogue System

Production RAG system for multi-turn clinical Q&A, grounded in 18K+ medical papers and guardrailed for patient safety. Built on top of 4 peer-reviewed healthcare AI publications.

18K+
Papers Indexed
<180ms
Time to First Token
94%
Intent Accuracy
RAGLangChainMedCPTMilvusFastAPIPydanticPython
Read the case study
03
03 / 05Featured
2024 – Present
GitHub

LUNA: Conversational AI Engine

Production-grade, fully local conversational AI assistant with multi-turn dialogue, an on-device voice pipeline, and persistent memory. Zero cloud. Zero subscriptions. Complete ownership.

8
LLM Backends
<50ms
UI Response
Conversation Memory
PythonFastAPILangChainChromaDBWhisperOllamaElectron
Read the case study
04
04 / 05
2026
GitHub

wLLM: Native Windows LLM Inference Server

Production-grade LLM inference server built from scratch for Windows. vLLM doesn't run natively there, so wLLM reimplements its core ideas — PagedAttention, continuous batching, CUDA graphs — as a native CUDA/C++ and PyTorch engine, no WSL2 or Docker required, with an OpenAI-compatible API on top.

3-7×
CUDA Graph Throughput
5
OpenAI-Compatible Endpoints
0
WSL2/Docker Required
PythonPyTorchCUDAFastAPIPydantic
Read the case study
05
05 / 05
Jan – May 2025

TITANS: Long-Context Dialogue Memory

Memory-augmented transformer architecture built to solve the core problem in conversational AI: retaining context across long multi-turn exchanges without quadratic attention cost.

+35%
Multi-turn Recall
−30%
Inference Latency
Throughput
PyTorchTritonvLLMTransformersQuantizationFastAPI
Read the case study