Gregory Chin
Production LLM systems for healthcare, fintech & consumer AI
8+ years shipping RAG pipelines, multi-agent orchestration, and real-time voice/video AI at scale — backed by hands-on ML engineering and a polyglot backend spanning Python, Go, Java, C#, and Rust.
- Years Building Production AI
- 8+
- Years Building Production AI
- Daily Video Events Processed
- 150k+
- Daily Video Events Processed
- Concurrent AI Agents Orchestrated
- 1,000s
- Concurrent AI Agents Orchestrated
- Voice AI Satisfaction Score
- 4.8/5
- Voice AI Satisfaction Score
About Me
AI Full-Stack Engineer with 8+ years building LLM-powered systems, real-time voice/video platforms, and high-scale backend AI infrastructure across healthcare, fintech, and consumer AI. Deep expertise in RAG pipelines, multi-agent orchestration, and HIPAA-compliant clinical AI, backed by hands-on ML engineering and applied data science.
LLM & RAG Systems
Dense, sparse, and hybrid retrieval; multi-agent orchestration with LangGraph; DSPy and structured prompt engineering across production pipelines.
Real-Time Voice & Video AI
LiveKit, Pipecat, and WebRTC pipelines handling turn-taking, interruptions, and sub-second latency.
Responsible AI & Evaluation
LLM-as-judge evaluation, output guardrails, and HIPAA-aligned RBAC/PHI controls — safety and compliance built into the pipeline, not bolted on.
ML Engineering & MLOps
LoRA/QLoRA fine-tuning, MLflow and Weights & Biases experiment tracking, reproducible model serving.
Polyglot Backend at Scale
Python for AI/ML, Go for high-throughput APIs, Java/Spring and C#/.NET for enterprise, Rust for performance-critical modules.
High-Scale Production Systems
150k+ daily video events, thousands of concurrent autonomous agents, and millions of daily AI-generated events sustained in production.
Experience
Senior AI Software Engineer
CurrentSelf-Employed
Jul 2025 – PresentVoice AI agents, multi-agent RAG orchestration, and enterprise LLM deployments for healthcare and SaaS clients — including Glass Health and Ringfree.
Senior Software Engineer
Butterflies AI
Dec 2023 – Jun 2025
Led a 7-person team building an AI social network on AWS Bedrock, serving millions of daily users.
Software Engineer
Theoria Medical
Feb 2019 – Nov 2023
Led a 4-person platform team building HIPAA/SOC2 clinical AI processing thousands of encounters monthly.
Software Engineer
Dayta AI
Aug 2017 – Jan 2019
Built real-time IoT pipelines for the Cyclops retail intelligence platform, Hong Kong.
B.S. Computer Science
Hong Kong College of Technology
2013 – 2017
Case Studies
Production RAG pipelines and multi-agent AI systems across healthcare, consumer AI, and retail intelligence — real problems, real stack decisions, measured impact.
Voice AI & Multi-Agent Healthcare Automation
Senior AI Software Engineer, Self-Employed · Jul 2025 – Present
Healthcare and SaaS clients — including Glass Health and Ringfree — needed to automate intake, insurance pre-auth, scheduling, and clinical documentation without sacrificing accuracy or compliance.
Why this stack: LiveKit + Pipecat for sub-second voice latency, LangGraph for stateful multi-agent orchestration with shared memory, Pinecone/Qdrant for compliant low-latency retrieval — chosen for production reliability under real patient/customer load, not novelty.
AI Social Network at Scale
Senior Software Engineer, Butterflies AI · Dec 2023 – Jun 2025
Butterflies AI needed to support millions of daily users across web, iOS, and React Native — with thousands of concurrent autonomous AI characters maintaining persona consistency and safety in real time.
Why this stack: LangGraph on Bedrock for durable agent state across millions of daily events, LLM-as-judge for scalable safety evaluation without a human-review bottleneck, Supabase for RLS-backed auth that shipped fast without sacrificing isolation.
HIPAA-Compliant Clinical AI Platform
Software Engineer, Theoria Medical · Feb 2019 – Nov 2023
Theoria Medical needed to process thousands of patient encounters monthly — intake, insurance documents, clinical notes, and inbound calls — under strict HIPAA and SOC2 requirements.
Why this stack: FastAPI for isolated PHI/RBAC boundaries, pgvector alongside Pinecone for compliant hybrid retrieval, ASP.NET Core to match existing EHR vendor tooling — prioritizing HIPAA compliance and integration fit over greenfield flexibility.
Retail Intelligence from Real-Time Video
Software Engineer, Dayta AI · Aug 2017 – Jan 2019
The Cyclops retail intelligence platform needed to turn raw RTSP camera and sensor data from malls and retail stores into actionable foot-traffic and demographic insights at scale.
Why this stack: OpenCV for real-time edge inference on live camera feeds, Rails/Sidekiq for fast multi-tenant iteration, managed MediaConvert rendering instead of custom infra — optimized for throughput at 150k+ daily videos on a lean team.
Example Projects
A sample of freelance and independent engagements across AI, real-time communication, and full-stack development.
Social AI

Healthcare AI



LiveKit Examples


AI Mobile Apps


Video Streaming



Other AI Tools


Skills & Toolset
A cross-stack toolkit spanning applied AI/ML, real-time systems, and polyglot backend engineering.
AI / LLM Engineering
Agentic Frameworks & Platforms
ML Engineering / MLOps
Vector Search & RAG Infra
Voice, Video & Real-Time
Backend by Language
Data & Cloud Infra
Data Science & Analytics
Let's Build Something Production-Grade
Open to full-time roles, consulting engagements, and collaborations in LLM systems, voice AI, and healthcare technology. Reach out or grab a copy of my resume.
Contact Information
Connect
View Profile
Open to full-time opportunities, consulting projects, and collaborations in Voice AI, real-time communication, healthcare technology, and LLM-powered SaaS.