AI Full-Stack Engineer

Gregory Chin

Production LLM systems for healthcare, fintech & consumer AI

8+ years shipping RAG pipelines, multi-agent orchestration, and real-time voice/video AI at scale — backed by hands-on ML engineering and a polyglot backend spanning Python, Go, Java, C#, and Rust.

RAG & RetrievalMulti-Agent / Agentic AILLM Fine-Tuning (LoRA/QLoRA)AI Safety & EvaluationVoice & Multimodal AI
Years Building Production AI
8+
Years Building Production AI
Daily Video Events Processed
150k+
Daily Video Events Processed
Concurrent AI Agents Orchestrated
1,000s
Concurrent AI Agents Orchestrated
Voice AI Satisfaction Score
4.8/5
Voice AI Satisfaction Score

About Me

AI Full-Stack Engineer with 8+ years building LLM-powered systems, real-time voice/video platforms, and high-scale backend AI infrastructure across healthcare, fintech, and consumer AI. Deep expertise in RAG pipelines, multi-agent orchestration, and HIPAA-compliant clinical AI, backed by hands-on ML engineering and applied data science.

LLM & RAG Systems

Dense, sparse, and hybrid retrieval; multi-agent orchestration with LangGraph; DSPy and structured prompt engineering across production pipelines.

Real-Time Voice & Video AI

LiveKit, Pipecat, and WebRTC pipelines handling turn-taking, interruptions, and sub-second latency.

Responsible AI & Evaluation

LLM-as-judge evaluation, output guardrails, and HIPAA-aligned RBAC/PHI controls — safety and compliance built into the pipeline, not bolted on.

ML Engineering & MLOps

LoRA/QLoRA fine-tuning, MLflow and Weights & Biases experiment tracking, reproducible model serving.

Polyglot Backend at Scale

Python for AI/ML, Go for high-throughput APIs, Java/Spring and C#/.NET for enterprise, Rust for performance-critical modules.

High-Scale Production Systems

150k+ daily video events, thousands of concurrent autonomous agents, and millions of daily AI-generated events sustained in production.

Experience

Senior AI Software Engineer

Current

Self-Employed

Jul 2025 – Present

Voice AI agents, multi-agent RAG orchestration, and enterprise LLM deployments for healthcare and SaaS clients — including Glass Health and Ringfree.

Senior Software Engineer

Butterflies AI

Dec 2023 – Jun 2025

Led a 7-person team building an AI social network on AWS Bedrock, serving millions of daily users.

Software Engineer

Theoria Medical

Feb 2019 – Nov 2023

Led a 4-person platform team building HIPAA/SOC2 clinical AI processing thousands of encounters monthly.

Software Engineer

Dayta AI

Aug 2017 – Jan 2019

Built real-time IoT pipelines for the Cyclops retail intelligence platform, Hong Kong.

B.S. Computer Science

Hong Kong College of Technology

2013 – 2017

Case Studies

Production RAG pipelines and multi-agent AI systems across healthcare, consumer AI, and retail intelligence — real problems, real stack decisions, measured impact.

Voice AI & Multi-Agent Healthcare Automation

Senior AI Software Engineer, Self-Employed · Jul 2025 – Present

View Client Product: Glass Health

Healthcare and SaaS clients — including Glass Health and Ringfree — needed to automate intake, insurance pre-auth, scheduling, and clinical documentation without sacrificing accuracy or compliance.

LiveKitLangGraphRAGPinecone/QdrantLoRA/QLoRARedisFastAPI

Why this stack: LiveKit + Pipecat for sub-second voice latency, LangGraph for stateful multi-agent orchestration with shared memory, Pinecone/Qdrant for compliant low-latency retrieval — chosen for production reliability under real patient/customer load, not novelty.

AI Social Network at Scale

Senior Software Engineer, Butterflies AI · Dec 2023 – Jun 2025

View Live Product: Butterflies AI

Butterflies AI needed to support millions of daily users across web, iOS, and React Native — with thousands of concurrent autonomous AI characters maintaining persona consistency and safety in real time.

AWS BedrockLangGraphStable DiffusionComfyUISupabaseReact Native

Why this stack: LangGraph on Bedrock for durable agent state across millions of daily events, LLM-as-judge for scalable safety evaluation without a human-review bottleneck, Supabase for RLS-backed auth that shipped fast without sacrificing isolation.

HIPAA-Compliant Clinical AI Platform

Software Engineer, Theoria Medical · Feb 2019 – Nov 2023

Theoria Medical needed to process thousands of patient encounters monthly — intake, insurance documents, clinical notes, and inbound calls — under strict HIPAA and SOC2 requirements.

HIPAA/SOC2TwiliopgvectorFastAPIASP.NET CoreRBAC

Why this stack: FastAPI for isolated PHI/RBAC boundaries, pgvector alongside Pinecone for compliant hybrid retrieval, ASP.NET Core to match existing EHR vendor tooling — prioritizing HIPAA compliance and integration fit over greenfield flexibility.

Retail Intelligence from Real-Time Video

Software Engineer, Dayta AI · Aug 2017 – Jan 2019

The Cyclops retail intelligence platform needed to turn raw RTSP camera and sensor data from malls and retail stores into actionable foot-traffic and demographic insights at scale.

OpenCVRTSPFFmpegAWS MediaConvertRailsSidekiq

Why this stack: OpenCV for real-time edge inference on live camera feeds, Rails/Sidekiq for fast multi-tenant iteration, managed MediaConvert rendering instead of custom infra — optimized for throughput at 150k+ daily videos on a lean team.

Example Projects

A sample of freelance and independent engagements across AI, real-time communication, and full-stack development.

Social AI

Butterflies AI
Butterflies AI
The first AI social network where humans and AI characters coexist. Led full-stack development for character generation using ComfyUI, Stable Diffusion, and custom training pipelines. Built cross-platform mobile experiences with React Native.
Social AIGenerative AIReact NativeComfyUI
View Project

Healthcare AI

Vocca AI
Vocca AI
Enterprise voice AI platform for healthcare workflows. Architected real-time voice agents handling patient intake, clinical documentation, and provider communication using ASR/TTS pipelines with HIPAA-compliant infrastructure.
Voice AIHealthcareHIPAAASR/TTS
View Project
Hippocratic AI
Hippocratic AI
Clinical AI automation platform providing decision support for medical professionals. Implemented generative AI for email drafting, clinical note summarization, and visual content generation for patient education.
HealthcareGenerative AIClinicalLLM
View Project
Glass Health
Glass Health
Medical knowledge AI empowering clinicians with instant clinical insights. Developed AI-driven diagnostic tools, integrated secure APIs for intake and summarization consumed by web apps and call workflows.
HealthcareMedical AIAPIsNLP
View Project

LiveKit Examples

Icon Streamer
Icon Streamer
Live streaming platform where voices unite and stories amplify. Built real-time meeting and streaming infrastructure using LiveKit, WebRTC, and scalable media servers for seamless multi-participant broadcasts.
LiveKitStreamingWebRTCReal-Time
View Project
AI Avatar + RAG + Pinecone
AI Avatar + RAG + Pinecone
Real-time AI avatar system with retrieval-augmented generation. Built using LiveKit for WebRTC video, Pinecone for vector search, and GPT-4 for context-aware responses with lifelike avatar rendering.
LiveKitRAGPineconeAvatarWebRTC
View Project
Call Transfer + Meeting AI Agent
Call Transfer + Meeting AI Agent
Intelligent call routing with LiveKit telephony integration (Twilio, FreeSWITCH). Built smart call transfer, automated transcription, real-time meeting summaries, and reduced latency across high-traffic environments.
LiveKitTwilioFreeSWITCHAI Agent
View Project
Glite AI
Glite AI
Next-gen AI communication platform on LiveKit. Developed real-time audio pipelines managing interruptions, silence detection, and turn-taking. Optimized jitter/packet loss for streaming reliability under concurrent load.
LiveKitVoice AIWebRTCReal-Time
View Project

AI Mobile Apps

The Pearl
The Pearl
iOS AI companion app delivering personalized conversational experiences. Built with React Native featuring intelligent chat, voice interactions, and contextual awareness powered by advanced NLP models.
iOSReact NativeNLPMobile
View Project
Magic AI Room
Magic AI Room
Android generative AI chat application powered by GPT models. Features image generation, text-to-speech, and multi-modal AI interactions with seamless cloud integration and offline capabilities.
AndroidGPTGenerative AIMobile
View Project

Video Streaming

Instantly Inc
Instantly Inc
Enterprise real-time video streaming platform. Architected scalable streaming with WebRTC, custom CDN infrastructure, adaptive bitrate, and ultra-low latency delivery for thousands of concurrent viewers.
StreamingWebRTCCDNReal-time
View Project
UHF App
UHF App
High-performance video streaming and content delivery platform. Implemented adaptive streaming protocols, content caching strategies, and global edge network integration for optimal playback worldwide.
StreamingVideoCDNEdge
View Project
Swip TV
Swip TV
Interactive video streaming with social features and real-time engagement. Built React-based video player with custom controls, live chat integration, and interactive overlays for enhanced viewer experiences.
StreamingInteractiveSocialReact
View Project

Other AI Tools

Creatify AI
Creatify AI
AI-powered video advertisement generation platform. Developed generative AI pipeline for script-to-video conversion, automated voiceovers, dynamic scene composition, and brand-consistent ad creation at scale.
AI VideoMarketingGenerative AIAutomation
View Project
Tagshop AI
Tagshop AI
AI-driven e-commerce solution with intelligent product tagging and visual search. Implemented computer vision models for automatic product recognition, similarity matching, and personalized shopping recommendations.
AIE-commerceComputer VisionML
View Project

Skills & Toolset

A cross-stack toolkit spanning applied AI/ML, real-time systems, and polyglot backend engineering.

AI / LLM Engineering

RAG (dense/sparse/hybrid)Multi-Agent OrchestrationLoRA/QLoRA Fine-TuningPrompt EngineeringDSPyGPT-4/5ClaudeGeminiMistral

Agentic Frameworks & Platforms

LangChainLangGraphLangSmithLangfuseAutoGenCrewAIMCPAWS BedrockAzure AI FoundryVertex AI

ML Engineering / MLOps

PyTorchscikit-learnXGBoost/LightGBMMLflowWeights & BiasesDVCFeastSageMakerONNXBentoML/Triton

Vector Search & RAG Infra

PineconeQdrantMilvusWeaviatepgvectorChromaOpenSearchElasticsearchHybrid SearchReranking

Voice, Video & Real-Time

LiveKitWebRTCPipecatTwilioDeepgramElevenLabsCartesiaWhisperFFmpegOpenCV

Backend by Language

Python / FastAPIGoJava / Spring BootC# / .NETRustTypeScript / NestJSRuby on Rails

Data & Cloud Infra

PostgreSQLRedisMongoDBSupabaseAWSGCPAzureDockerKubernetesTerraform

Data Science & Analytics

Pandas / PolarsSQLA/B TestingSpark / PySparkTime-Series AnalysisComputer VisionNLP

Let's Build Something Production-Grade

Open to full-time roles, consulting engagements, and collaborations in LLM systems, voice AI, and healthcare technology. Reach out or grab a copy of my resume.

Contact Information

Connect

LinkedIn

View Profile

Open to full-time opportunities, consulting projects, and collaborations in Voice AI, real-time communication, healthcare technology, and LLM-powered SaaS.