Singapore · Sovereign AI & Private LLM Specialist · Available for Roles
Kiran Machha

Senior AI Engineer

Forward Deployed AI Infrastructure Specialist

Taking AI solutions from business problems to production-grade systems inside client VPCs, air-gapped networks, and Kubernetes GPU clusters.

11 Real Systems

Local Codebase Repos

100% Private

Air-Gapped & On-Prem

vLLM & Rust

GPU Inference Engine

Full Lifecycle

Problem → Cost Opt

End-To-End Architecture Flow

The Complete Enterprise AI Pipeline

Production Pipeline Nodes
01

Data Ingestion

OCR, Parsing, Chunking, Embeddings

02

Model Engine

Open-Weight LLMs, vLLM, Fine-Tuned LoRA

03

AI Application

Agents, RAG, Typed Tool Calling

04

APIs & Gateways

FastAPI, Semantic Router, Rate Limiter

05

Enterprise Systems

Core Banking, ERP, CRM, Identity

06

Kubernetes Cluster

GPU Autoscaling, vLLM Node Pools

07

Cloud & On-Prem

Private VPC, Air-Gapped Datacenters

08

Monitoring & Tracing

OpenTelemetry, Self-Hosted Evals

09

Alerts & Governance

Hallucination Checks, Drift Detection

10

Cost Optimization

Semantic Caching, Quantization

Kiran Machha

Kiran Machha

Senior AI Engineer | Forward Deployed Engineer

Singapore · Founder @ KaviAI

Forward Deployed Engineering

Embedding inside customer teams to ship Sovereign AI that works inside their walls.

As a Senior AI Engineer & Forward Deployed Engineer (FDE), I sit inside customer engineering repos, infrastructure clusters, and compliance reviews to build, deploy, and scale AI systems.

My specialism is Sovereign AI & Private On-Premises Infrastructure — tuning and serving open-weight models (Llama 3, Qwen2, DeepSeek, Mistral) in private VPCs, air-gapped data centers, or local GPU nodes for banks, healthcare systems, and regulated enterprises.

From fine-tuning open weights to containerized vLLM serving, hybrid vector retrieval, agent tool calling, and Kubernetes GPU scaling — I deliver fully operational AI systems with complete cost optimization and self-hosted observability.

Core Expertise
Private AI & On-Premises AI
Machine Learning & Deep Learning
LLM Application Development
LLM Training, Fine-tuning, Deployment & Management
RAG & Vector Search
AI Agents & Agentic Workflows
OCR, NLP, Computer Vision & Voice AI
Prediction & Forecasting
Kubernetes & GPU Infrastructure
AI/LLM Observability, Monitoring & Alerting
AWS, Azure & Google Cloud AI
AI Infrastructure, Architecture & Cloud Cost Estimation
Real System Repositories

11 Featured Production Systems

Click any project card to open an interactive deep dive into its architecture flow, API endpoints, performance metrics, and tech stack.

01Private Local AI Platform

ODS — Osmantic Deployment System

Turn your PC, Mac, or Linux box into a self-hosted private AI server with a single command

A self-hosted AI deployment platform built around 24 bundled Docker service manifests, hardware-accelerated overlays (NVIDIA, AMD, Apple Silicon, Intel Arc), a control dashboard, LiteLLM gateway, RAG pipeline, local voice STT/TTS, and privacy tools.

Single-command bootstrap (`curl -fsSL install.osmantic.com/ods.sh | bash`) for Linux, macOS, and Windows WSL2
Bundles 24 service manifests: llama-server, Open WebUI, LiteLLM, Qdrant, ComfyUI, Whisper, Kokoro, n8n, SearXNG
Hardware auto-detection applying compose overlays for NVIDIA CUDA, Apple Metal, AMD ROCm, and Intel Arc GPUs
Shell Bootstrap & PowerShell CLIllama-server & OllamaLiteLLM Proxy (:4000)Open WebUI (:3000)ODS Dashboard (:3001)+3
02Autonomous Local Agent

KaviAgent — Local-First Personal AI Assistant

Local-first personal assistant with SQLite memory state, local web cockpit, and 95-line plain Python loop

A local-first personal AI assistant framework demonstrating the four pillars of serious agents: Harness, Reasoning Loop, Stateful Memory (SQLite), and LLM-as-Judge Evals. Features a local web dashboard cockpit at localhost:7777.

Local-first memory architecture stored in a single SQLite database file (`.kavi/state.db`)
Three-tier memory system: semantic facts, episodic conversations, and procedural skills
Smart retrieval gate evaluating per turn whether memory retrieval is required
Python 3.12 (UV package manager)SQLite (.kavi/state.db)FastAPI & Static Web Cockpit (:7777)Telegram Bot APIClaude, OpenAI, DeepSeek, Ollama
03Desktop Multi-Agent IDE

KaviSpace — Tauri v2 Multi-Agent Desktop & Swarm Platform

Desktop application with Tauri v2, xterm.js terminals, and KaviSwarm multi-agent pipeline

A Tauri v2 + Rust desktop application and multi-agent development environment that autonomously writes, builds, deploys, and live-demos full-stack applications through a three-phase AI swarm pipeline (KaviSwarm).

Tauri v2 + Rust native desktop shell with React 19 and Vite frontend
Integrated xterm.js terminal emulation powered by node-pty and WebSockets
KaviSwarm 3-phase autonomous pipeline: PRD generation, build execution, and Playwright verification
Tauri v2 (Rust)React 19, TypeScript, Vite & Zustandxterm.js & node-ptyNode.js / Express (:3001) & WebSockets@anthropic-ai/sdk (Claude API)+1
04Multi-Channel AI SaaS

KaviAI — Growth & Multi-Channel Content Platform

AI agent platform planning, generating, and publishing content across 10+ platforms automatically

An enterprise AI agent platform for growth, marketing, and distribution. Uses Google Gemini with LiteLLM gateway and OpenAI fallback to automate SEO, Reddit, LinkedIn, X, Instagram, and YouTube publishing.

Multi-channel publishing pipeline across Reddit, LinkedIn, X, Instagram, Facebook, Threads, YouTube, and TikTok
Powered by Google Gemini 2.5 Flash with LiteLLM proxy and GPT-4o-mini fallback
Supabase backend with PostgreSQL Row-Level Security (RLS) and SSR cookie authentication
Next.js 16 App Router & React 19Vercel AI SDK (ai 6.x)Google Gemini 2.5 FlashSupabase (PostgreSQL + RLS)Tailwind CSS v4
05Developer AI Infrastructure

Agentic Coding Platform & Workspaces

Self-hosted AI development infrastructure for secure, governed agentic coding

A self-hosted developer platform providing containerized workspace environments, AI coding agents, and governance controls. Allows developers and AI agents to code side-by-side inside controlled sandbox environments.

Container-based sandbox isolation with explicit CPU/RAM limits per developer workspace
Multi-agent runner orchestrating autonomous coding tasks across repository AST trees
Model Context Protocol (MCP) tool integration with Cursor, Claude Code, and Windsurf
FastAPIDocker Engine & Kubernetes APIModel Context Protocol (MCP)UV (Python) & Pnpm
06Enterprise Hybrid RAG

Corporate Organization RAG System

Production-grade Retrieval-Augmented Generation for enterprise knowledge management

A complete corporate organization RAG system that ingests internal documents, enables hybrid search across organizational knowledge, and provides intelligent Q&A through agentic retrieval with LangGraph.

Automated document ingestion & PDF chunking pipeline via Airflow 3.0
Hybrid Reciprocal Rank Fusion (RRF) matching BM25 keyword precision with vector semantics
Agentic LangGraph workflow featuring document grading & automatic query rewriting
OpenSearch 2.19 (Hybrid BM25 + Vector)FastAPI 0.115+LangGraph & LangChainLocal Ollama LLM
07Banking & Financial AI

Bank Cheque OCR Automation

Handwriting-aware cheque processing with automatic fraud detection for regulated banking

A production-ready Bank Cheque OCR API automating the extraction and verification of critical information from scanned bank cheques. Designed for high-volume banking back-offices processing 5,000+ cheques daily.

Sub-second field extraction: Payer, Payee, Numeric Amount, Written Amount, Cheque Number, IFSC, MICR data
Dual-support for both printed bank text and complex handwritten entries
Automatic cross-validation comparing numeric figures vs written English text
FastAPIOpenCV & PyTesseractPydantic v2
08Realtime Voice AI

Realtime Customer Service Voice Agent

Production-ready realtime voice agent for customer service call centers

A complete customer service voice agent system capable of handling inbound and outbound telephone calls with sub-second speech recognition, intelligent tool-calling responses, knowledge lookup, and conversation tracing.

Inbound and outbound telephone call handling via Twilio Webhook integration
Ultra-low-latency realtime conversational audio loop powered by FastRTC
Multi-avatar persona system supporting distinct department voices and personalities
FastRTCTwilio WebRTC / SIPSuperlinked & Qdrant
09Agentic Workflows

n8n AI Agents & Workflow Automation Collection

Production-ready n8n automation workflows powered by AI agents for enterprise operations

A curated collection of 17 enterprise-grade n8n automation workflows integrating LLM agents, automated triage, email notification generators, IT ticket processors, and daily reporting systems.

17 production workflows covering IT, DevOps, Customer Support, HR, and Sales
Daily Server & API Health Monitor sending automated email alerts
Customer Support Auto-Responder generating contextual AI draft responses
n8n Workflow AutomationOpenAI (GPT-4o / GPT-4o-mini)Gmail SMTP & Webhooks
10Sovereign GPU Pipeline

Invoice OCR & Multi-Scale Processing Pipeline

High-throughput document parsing with vLLM, Rust API gateway, and async GPU queues

A multi-stage asynchronous invoice processing system engineered with vLLM vision model inference, a high-concurrency Rust API gateway, and async task queues for enterprise accounting teams.

Rust-based gateway capable of receiving thousands of concurrent document uploads
vLLM vision model inference serving Qwen2-VL and Donut models on GPU instances
Automatic line-item extraction, subtotal/total reconciliation, and tax/VAT calculation
Rust (Axum & Tokio)vLLM Vision EngineRedis & Celery
11Developer AI Tooling

PageBolt MCP Server for AI Coding Assistants

Model Context Protocol (MCP) server giving AI agents web capture, screenshots, and page inspection

An open-source Model Context Protocol (MCP) server connecting AI coding assistants (Cursor, Windsurf, Claude Desktop, Cline) to PageBolt capture APIs for screenshotting, PDF generation, and page inspection.

9 specialized MCP tools: take_screenshot, generate_pdf, create_og_image, inspect_page, observe_page, record_video
Token-budgeted page observation specifically optimized for AI browser agents
Device preset support covering 25+ viewports (iPhone, iPad, MacBook, Galaxy)
Model Context Protocol (MCP)TypeScript & Node.jsCursor, Claude Desktop, Windsurf, Cline
Complete Engineering Lifecycle

Business Problem → Production AI

Not just an AI developer. An FDE who owns the full lifecycle from initial business discovery to GPU cost optimization.

01

Business Problem

Identify workflow bottlenecks, regulatory constraints, ROI targets, and latency requirements.

02

AI Architecture

Design model strategy, open vs closed weights, vector storage, context windows, and safety barriers.

03

Development

Fine-tune open-weight models (LoRA/QLoRA), build RAG pipelines, typed agent tools, and evaluation harnesses.

04

Enterprise Integration

Connect models to ERPs, core banking APIs, CRM database queues, and OAuth/RBAC identity systems.

05

Deployment

Containerize with vLLM/Triton, deploy on Kubernetes, GPU autoscaling, and zero-downtime rollouts.

06

Observability

Self-hosted tracing, evals, hallucination monitoring, semantic logging, and token usage analytics.

07

Production Operations

Automated failover, load balancing, continuous eval gates, and human-in-the-loop fallback queues.

08

Cost Optimization

Semantic caching, model distillation, dynamic batching, and GPU node pool scale-to-zero.

Integration Capabilities

Integrating AI Into Operational Systems

Models add value when integrated cleanly into business workflows, databases, and core APIs.

RAG

Integrated

Vector Search

Integrated

OCR

Integrated

NLP

Integrated

Vision

Integrated

Voice

Integrated

AI Agents

Integrated

Prediction

Integrated

Forecasting

Integrated

Automation

Integrated
Technology Stack

Open-Source & Self-Hosted Stack

Proven technology choices enabling enterprises to run AI models on-premises or in private clouds without third-party API dependencies.

AI/ML

11 Tools
PyTorchHugging FaceTransformersMLDeep LearningNLPOCRVisionVoicePredictionForecasting

LLM

10 Tools
RAGEmbeddingsVector SearchRerankingFine-TuningLoRA/QLoRALLM TrainingLLM DeploymentModel ManagementEvaluation

Agents

6 Tools
AI AgentsAgentic AIMCPTool CallingMulti-Agent Workflowsn8n

Infrastructure

7 Tools
DockerKubernetesGPUvLLMTritonModel ServingCI/CD

Data

5 Tools
PostgreSQLpgvectorOpenSearchElasticsearchVector Databases

Cloud

7 Tools
AWSAzureGoogle CloudCloud AIGPU InfrastructureCost EstimationCost Optimization

Observability

8 Tools
LogsMetricsTracingLLM ObservabilityModel MonitoringPerformance MonitoringCost MonitoringAlerts
Career Track Record

Enterprise Delivery Experience

From regulated banking change control to SaaS GPU platforms and forward-deployed client engagements.

Founder & Forward Deployed AI Engineer

KaviAI · kaviagentic.com

PresentSingapore
  • Builds and deploys agentic AI systems with customers — multi-agent orchestration, tool calling, and self-correcting reasoning loops that survive real inputs.
  • Runs the full LLMOps chain: serving, gateway routing, evaluation harnesses and observability, so quality regressions are caught before customers find them.
  • Ships production infrastructure on Kubernetes with GPU autoscaling, and publishes the patterns as open workshops.

AI Platform / LLM Ops Engineer

Enterprise SaaS

PriorSaaS
  • Stood up the shared inference platform — model gateway, quota and usage-based billing, and semantic caching to hold cost per request down as traffic grew.
  • Introduced evaluation gates into CI so prompt and model changes shipped on evidence rather than vibes.
  • Moved batch and streaming inference onto autoscaling GPU node pools with scale-to-zero for off-peak hours.

Full Stack AI Engineer

Payments

PriorPayments
  • Delivered customer-facing AI features end to end — Next.js and FastAPI through to the retrieval and ranking layers behind them.
  • Built hybrid search and RAG over transaction and policy corpora, with citation enforcement for audit trails.
  • Hardened the path to production: async workers, idempotent retries and structured tracing across services.

Software Engineer — Enterprise Applications

Banking

EarlierBanking
  • Built and maintained enterprise-grade applications under regulated change control, where a failed release is an incident report.
  • Microservices and API platforms on Spring Boot and Node.js, containerised and delivered through automated pipelines.
  • The grounding that makes the AI work deployable: security review, access control and data handling as defaults rather than afterthoughts.
Get In Touch

Need a Senior AI Engineer or FDE?

Available for Senior AI Engineer, Forward Deployed Engineer (FDE), or Sovereign AI Infrastructure contracts and permanent roles across Singapore, APAC, and EMEA.

Direct Email

machhakiran@gmail.com

Primary contact method.

LinkedIn

/in/machhakiran

Full career history.

GitHub

machhakiran

Agents, RAG & GPU infra.

Company

KaviAI →

Agentic workshops & platform.