Inside Lila (Project JARVIS): How Rishabh Joshi Built a 10-Subsystem Autonomous AI Operating System
Rishabh Kanhaiya Joshi
Founder, Treliotech & Creator of Lila
Key Takeaways & Executive Summary
- Architected by Rishabh Joshi (VIT Pune Computer Engineering) as an embodied 10-subsystem autonomous AI operating system.
- Dual-speed cognitive hierarchy: sub-400ms reactive conversational loop coupled with a deliberative SQLite WAL cortex.
- Level 5+ Computer Use with 1000x1000 sub-pixel scaling, 3-gate truthful checklist progression, and Python-level Hard Blocker Gates.
- Air-gapped near-ultrasound acoustic telepathy modem (15.5 kHz - 18.2 kHz) for zero-network cross-device pairing.
- Audited benchmark performance suite showing real, measured numbers across response latency, turn reduction, and memory footprint.

The Origin & Core Philosophy: Why Rishabh Joshi Built Lila (Project JARVIS)
While studying Computer Engineering in his second year at Vishwakarma Institute of Technology (VIT) in Pune, Treliotech founder Rishabh Kanhaiya Joshi observed a fundamental limitation in the current generation of artificial intelligence tools: conversational chatbots live in an API vacuum, completely isolated from the operating system, native desktop software, and real user workflows.
Traditional assistants like Siri, Alexa, or standard chat interfaces only answer questions inside a text box. When asked to actually perform work—such as filling out a travel booking, auditing a multi-sheet spreadsheet, or adjusting complex operating system settings—they demand that the human operator copy-paste code, manually click buttons, or repeatedly answer annoying confirmation prompts.
To bridge the gap between passive language models and embodied desktop autonomy, Rishabh architected Lila (codenamed Project JARVIS in the codebase). Built directly on a Windows 11 host with Linux VM benchmark parity, Lila operates under two core engineering doctrines: the Zero-Confirmation Barrier (bypassing unnecessary defensive prompts while halting strictly at financial payment gates) and the Zero-Hardcoding Doctrine (eliminating brittle CSS selectors, static domain lists, and hardcoded coordinates in favor of universal multimodal perception).
The Zero-Hardcoding Doctrine: Lila contains zero static domain dictionaries, zero rigid XPath selectors, and zero hardcoded pixel tables. All destinations, elements, and coordinates are resolved dynamically through real-time vision, accessibility trees, and semantic reasoning.
High-Level Runtime Topology & StateBridge (:8765)
Lila does not run as a monolithic executable. She is structured as a coordinated, distributed topology of local processes communicating over high-speed loopback channels on the host machine. At the center of this topology is StateBridge (core/state_bridge.py), an asynchronous WebSocket hub running on ws://127.0.0.1:8765.
StateBridge coordinates the Python 3.11 intelligence core, the Brave Browser Manifest V3 extension, the multi-window transparent Electron overlay, and native Windows Win32 hardware actuation drivers. To connect synchronous agent reasoning loops with asynchronous network I/O, StateBridge implements a Correlated Asynchronous RPC engine: every outbound command is stamped with a UUIDv4 future identifier, allowing Python worker threads to await browser DOM actions without freezing the master event loop.
- Continuous State Broadcasts: Streams mood shifts, audio RMS levels, live captions, and task checklist arrays at 30–60 FPS to the Electron presentation HUD.
- Cryptographic Pairing Handshake: Secures the browser extension via 32-hex pairing tokens with zero-touch localhost auto-repair and exponential reconnection backoff.
- 300ms Trailing Selection Debounce: Ingests user mouse text selections from active web pages without flooding the Python backend with microsecond drag events.
Dual-Mode Actuation: In-Browser Copilot vs. OS Computer Use
A key breakthrough engineered by Rishabh Joshi is Lila’s Dual-Mode Actuation Engine. Rather than forcing all actions through slow visual screenshot scraping or relying solely on brittle browser drivers, Lila dynamically switches between two complementary execution modes based on the active application.
In Mode A (In-Browser Copilot), Lila interacts directly with web pages inside the user’s everyday Brave browser profile. She traverses Shadow DOM trees, computes accessible element names according to W3C ARIA standards, and bypasses modern React 16/17/18/19 synthetic event tracking by invoking native HTMLInputElement prototype descriptors and resetting _valueTracker. In Mode B (OS Computer Use), Lila captures high-resolution screenshots, maps targets to a standardized 1000x1000 virtual sub-pixel grid, and actuates any native desktop app—from VS Code and LibreOffice to GIMP and Windows Settings.
Modern web forms often clear scripted text on blur. Lila invokes Object.getOwnPropertyDescriptor(HTMLInputElement.prototype, "value").set directly and clears el._valueTracker, ensuring 100% field persistence on platforms like RedBus, MakeMyTrip, and Flipkart.
Ten Breakthrough Level 5+ Autonomous Upgrades (U1 – U10)
In standard Level 5 automation, agents often get stuck in operational traps: search-query hopping after looking only at sponsored ads, misclick deadlocks caused by dynamic ad banners, or infinite modal open/close loops. To eliminate these failure modes, Rishabh engineered ten interlocking autonomous resilience upgrades (U1 through U10):
- U1 (Adaptive Scroll Threshold): Calculates horizontal scanline brightness density variance to require 2 to 6 deep scrolls past sponsored ad carousels before allowing query reformulations.
- U2 (Structured Page State Injection): Telemeters active URL, document title, and live vertical scroll percentage ([PAGE STATE] Scroll: 42.5%) directly into the model prompt to eliminate spatial disorientation.
- U3 (Smart 8-Entry Action History): Tracks an 8-step memory buffer tagged with verified pixel diff outcomes (✅ / ❌) and injects [STUCK WARNING] banners on streaks >= 1.
- U4 (Gemini Vision Re-Grounding): Spawns a lightweight gemini-2.5-flash-lite sub-agent on misclicks (~200 tokens, ~300ms) to re-localize displaced buttons without failing the mission.
- U5 (Universal LLM Site Resolver): Dynamically discovers canonical URLs for any natural language domain using an LLM extractor backed by an in-memory MD5 cache.
- U6 (Keyboard-First Fallback): Synthesizes accessibility standard Tab + Return keystrokes when clicks on custom div hitboxes fail after two consecutive turns.
- U7 (Dynamic Tab Entropy Cleaner): Automatically terminates rogue advertising and popup tabs using heuristic title signals ("Win", "Spin", "Gift Card") while protecting a domain whitelist.
- U8 (Action Fingerprint Loop Detector): Hashes actions into a rolling 6-step MD5 ring buffer, mathematically breaking cyclic execution loops when an identical action repeats 3+ times.
- U9 (Per-Step Retry Budget): Wraps hardware actuation in an automatic 2-attempt retry loop with a 300ms OS reaction pause, swallowing 90% of transient system glitches.
- U10 (Parallel Screenshot Pipeline): Uses an asynchronous background thread executor to capture and downsample the next screen frame during verification, reducing turn latency by 30%.
Universal ConditionGuard & The Hard Blocker Gate
When an autonomous agent shops or books services on the open web, e-commerce dark patterns pose severe risks: product pages often default to smaller, cheaper variants (e.g. 200ml instead of 1000ml), or keep "Buy Now" active even when the requested size is out of stock.
To protect the user from costly purchasing mistakes, Rishabh designed Universal ConditionGuard (core/gemini_computer_use_agent.py). ConditionGuard operates in three strict phases: pre-flight JSON contract extraction, real-time DOM variant verification (inspecting variant pills, selection flags, and disabled/dotted CSS borders), and the Python-Level Hard Blocker Gate.
If an e-commerce page defaults to 200ml when the user requested 1000ml, ConditionGuard activates a physical Python firewall that blocks "Buy Now" and "Add to Cart" clicks, forcing the agent to select the correct variant pill before authorization.
Anti-Bot Ghost Cursor Dynamics: Minimum-Jerk Biomechanical Kinematics
Traditional robotic process automation (RPA) tools teleport the mouse cursor instantaneously in 0 milliseconds along straight Euclidean lines. Modern fraud and bot-detection engines (Cloudflare Turnstile, Akamai, Google reCAPTCHA) monitor mouse kinematics and flag zero-jerk, non-human jumps immediately.
Lila bridges the physical actuation gap with Ghost Cursor (windows_lila/ghost_cursor.py). Instead of teleporting, Lila calculates minimum-jerk polynomial velocity profiles based on human motor control research: s(u) = 10u^3 - 15u^4 + 6u^5. She injects a perpendicular biomechanical wrist arc proportional to distance and samples waypoints along a quadratic Bézier curve at 60 FPS.
The cursor renders inside a transparent Win32 canvas with WS_EX_LAYERED, WS_EX_TRANSPARENT, and WS_EX_TOPMOST styles. Because the window is click-through, the operating system routes all physical human mouse events directly to underlying software without interference.
Visual Presentation: 3D VRM Humanoid Companion & Floating HUD

Lila’s visual presentation layer decomposes desktop interaction into specialized, hardware-accelerated transparent surfaces managed by Electron (lila-overlay/main.js) running with Direct3D 11 and occlusion-disabling flags:
- 3D VRM Humanoid Avatar (three_avatar.js): Anchored at the bottom-right of the screen, Lila features a high-fidelity Three.js VRM companion with an eyelid clamp ceiling (Whappy <= 0.25) to preserve bright eye contact, bilateral synchronous blinks, and desktop-wide 30 FPS mouse parallax tracking.
- Phonetic Viseme Engine: Maps speech text tokens and audio RMS power to five distinct mouth visemes (aa, oh, ou, ee, ih) at natural conversational cadence (13.3 characters/second).
- Floating Task Progress HUD (taskProgressWin): Anchored at (28, 120), this Samsung AI-style glassmorphism checklist scales dynamically (Height = 130 + 44 * Steps) to display real-time step rings, glowing spinners, and verified green checkmarks.
- Intelligence Terminal (cardsOverlayWin): A draggable glassmorphism canvas rendering research dossiers, offline KaTeX mathematical equations, and side-by-side comparative views.
Cognitive Cortex, Memory Systems & Mission DAG Orchestration
To achieve sub-second reactivity while retaining long-term episodic memory, Lila divides cognition into a Dual-Speed Architecture. The Fast Reactive Path (fast_agent.py) delivers sub-400ms end-to-end voice and tool interaction using Gemini 2.5 Flash Native Audio and an in-memory prompt cache (<0.5ms retrieval). Turn logging is completely asynchronous: interaction payloads are pushed onto an in-memory queue in under 0.05ms, leaving the foreground conversational thread completely unblocked.
In the background, the Deliberative Cognitive Cortex (core/lila_cognitive_cortex.py) persists interactions into a unified SQLite Write-Ahead Logging (WAL) database (data/lila_memory.db). It indexes facts using native SQLite FTS5 full-text search, resolves relative temporal queries ("kal", "parso", "yesterday"), and automatically synthesizes first-person Daily Relationship Diaries (data/diary/YYYY-MM-DD.md) during idle periods.
Complex multi-tool tasks are orchestrated by the Level 5 Branching Mission DAG Engine (core/mission_dag.py). The DAG engine supports recursive {{node_id.key}} placeholder interpolation, automated fallback routing on tool failures, and an anti-cycle circuit breaker enforcing a hard ceiling of 15 steps.
Air-Gapped Acoustic Telepathy: The Ultrasonic Beam Modem
In restricted or air-gapped environments—where laptops and smartphones cannot connect to shared Wi-Fi or Bluetooth networks due to enterprise firewalls or security isolation—Lila transmits data through room air using the Ultrasonic Telepathy Beam (core/ultrasonic_beam.py).
Operating on a Dual-Tone Multi-Frequency (DTMF) modulation scheme across near-ultrasound frequencies (15.5 kHz to 18.2 kHz), the transmitter emits inaudible sound waves through standard laptop speakers at 11.1 baud. Symbols consist of a 65ms dual-tone followed by a 25ms silence guard interval, eliminating acoustic room reverberation. On the receiving smartphone, a lightweight Progressive Web App (web/ultrasonic_portal.html) running a 1024-point Web Audio API FFT demodulates the signal, verifies a modulo-256 checksum, and copies the token directly to the mobile clipboard.
Grounded Benchmark Suite: Real Measured Performance Metrics
Rather than relying on inflated synthetic benchmarks or marketing exaggerations, Lila is evaluated against strict, repeatable performance metrics on her production Windows 11 host and the Linux VM OSWorld benchmark suite. Every number in the table below represents audited, measured reality across real hardware and network environments:
These benchmark figures are audited directly from the production repository (jarvis_project). They represent true measured operational latency, mathematical error rejection, and proven hardware execution rather than synthetic theoretical claims.
| Operational Dimension | Traditional Agent / Naive RPA | Lila 2.0 (Project JARVIS) Benchmark | Audited Engineering Mechanism |
|---|---|---|---|
| Voice Dialogue Latency | 2.5s - 6.0s (STT -> LLM -> TTS -> MP3 file) | < 400 ms (Conversational Turn) | Gemini 2.5 Flash Native Audio + In-Memory 24kHz PCM PyAudio (~42ms chunks) |
| Prompt Memory Injection | 120ms - 450ms (Disk read / RAG round-trip) | < 0.5 ms (RAM Cache Retrieval) | Pre-computed in-memory prompt cache with 10s TTL & event-driven invalidation |
| Turn Logging Overhead | 40ms - 120ms (Synchronous SQLite/Disk I/O) | < 0.05 ms (Immediate Thread Return) | Non-blocking _episode_queue drained by LilaEpisodeWriter background daemon |
| Desktop Spreadsheet / Doc Task | 45s - 90s (12-15 fragile visual GUI steps) | < 250 ms (1 Single Turn) | Code-First Strategy Router + Headless openpyxl / python-docx execution |
| Windows Settings / App Launch | 3.5s - 7.0s (Start menu search + click lag) | < 100 ms (Instant Launch) | Windows URI Protocol Accelerator (ms-settings:*, calc, code, explorer) |
| Multi-Field Web Form Submission | 6.0s - 9.0s (3 separate visual turns) | < 800 ms (1 Single Turn) | Composite Action Chaining (composite_action) with 300ms DOM pauses |
| Mouse Trajectory Bot Detection | 100% Flagged (0ms robotic linear jumps) | 0% Flagged (Human Biomechanics Pass) | Minimum-Jerk polynomial s(u) = 10u^3 - 15u^4 + 6u^5 + Biomechanical Bézier Arc |
| E-Commerce Variant Compliance | Frequently Fails (buys 200ml default / OOS) | 100% Contract Compliance | Universal ConditionGuard + Python-level Hard Blocker Gate on Buy Now |
| Screen Capture Ingestion Latency | 120ms - 250ms (Disk write PNG -> read -> upload) | < 20 ms (Pure RAM Pipeline) | In-Memory io.BytesIO Quality 85 JPEG compression + Lanczos downsampling |
| Air-Gapped Telepathy Transmission | Impossible without Wi-Fi / Bluetooth / Cables | 1.8s (16-Nibble Token Transmission) | Near-Ultrasound DTMF Guard-Gap Modem (15.5-18.2 kHz, 11.1 baud) |
| API Quota Resilience (HTTP 429) | Fails completely / freezes user session | 0 Errors Surface to User | 48 Execution Channels (3 Keys x 16 Models) with instant 65s cooldowns |
| Automated Test Suite Execution | Slow / Flaky / Network Dependent | 3.12 seconds (40 Tests Passed) | Isolated __new__() mock harness, 0 network/display dependencies, 100% pass |
| Disaster Recovery Readiness | Manual backups / divergence prone | 100% Bit-for-Bit SHA-256 Parity | Automated dual-drive mirror across 72 files between C: and D: drives |
System Resilience, Quota Circuit Breakers & The Dual-Drive Parity Rule
Production computer-use agents must withstand severe operational stress: API rate-limiting spikes, transient network shed, and local hardware failures. Lila guards against cloud quota exhaustion through the Multi-Key Quota Circuit Breaker (core/quota_circuit_breaker.py). By rotating across three Google GenAI keys and a 16-model cascade spanning 6 operational tiers, Lila maintains 48 independent execution channels. If an HTTP 429 occurs, that candidate trips into an instant 65-second cooldown and Lila fails over to a healthy channel with zero perceptible user delay.
To guarantee complete regression safety, Lila’s codebase is protected by a 40-test automated verification suite (tests/) covering ConditionGuard contracts, Level 5+ upgrades, and truthful checklist reporting. The entire suite executes in 3.12 seconds with zero network roundtrips.
Finally, Lila enforces an absolute Dual-Drive SHA-256 Parity Rule: all 72 production modules, test suites, and documentation files are cryptographically mirrored between the primary C: workspace and a hot-standby D: drive using 64KB chunked hashing. In the event of a catastrophic solid-state drive failure, the entire operating system agent can boot from D:jarvis_project in under 60 seconds with zero data loss.
Frequently Asked Questions
Who created Lila and what is Project JARVIS?
Lila (Project JARVIS) was designed and built by Rishabh Kanhaiya Joshi, a 2nd-year B.Tech Computer Engineering student at Vishwakarma Institute of Technology (VIT Pune) and the founder of Treliotech. It is an embodied, Level 5 autonomous AI operating system for Windows 11 with comprehensive browser copilot and desktop computer-use capabilities.
What makes Lila different from cloud chatbots like ChatGPT or Claude?
Unlike traditional chatbots that live inside a browser text window and only produce text, Lila lives directly on your workstation. She sees your screen via continuous vision, executes code-first desktop workflows, controls apps with humanized anti-bot mouse trajectories, fills web forms while bypassing React synthetic event tracking, and communicates via a 3D VRM humanoid avatar and floating heads-up display.
What are the real benchmark response times for Lila?
Lila achieves < 400ms conversational voice latency using Gemini 2.5 Flash Native Audio and in-memory 24kHz PCM streaming. In-memory memory prompt injections take < 0.5ms, asynchronous turn logging takes < 0.05ms, headless spreadsheet tasks complete in < 250ms, and Windows app launches via protocol accelerators take < 100ms.
How does Lila bypass bot detection during browser automation?
Lila runs directly inside the user’s everyday Brave browser profile, inheriting existing logins, active sessions, and cookies without launching suspicious headless browsers (navigator.webdriver remains false). For physical mouse actuation, she uses Flash & Hogan minimum-jerk polynomial curves with biomechanical wrist arcs rather than robotic 0ms linear jumps.
How does Universal ConditionGuard prevent wrong product purchases?
ConditionGuard parses user constraints into pre-flight JSON contracts, audits real-time DOM variant pills for stock and selection state, and activates a Python-level Hard Blocker Gate that physically blocks clicks on "Buy Now" or "Add to Cart" if the active variant or price violates the user’s contract.
What is the Ultrasonic Telepathy Beam in Lila?
The Ultrasonic Beam is an air-gapped acoustic modem that transmits arbitrary text, 2FA tokens, and URLs over near-ultrasound sound waves (15.5 kHz to 18.2 kHz) from laptop speakers to a smartphone microphone via a mobile PWA portal, requiring zero Wi-Fi, Bluetooth, or cables.
How does Project JARVIS relate to Treliotech?
Project JARVIS (Lila) represents the core autonomous systems research and deep engineering foundation built by Rishabh Joshi. The agentic architectures, state machine resilience, and zero-disruption principles proven in Lila directly power Treliotech’s commercial 24/7 AI workflow and WhatsApp automation systems for clinics, enterprises, and high-growth practices.
Automate Your Practice WhatsApp in Under 24 Hours
Test 24/7 autonomous scheduling and triage on your existing WhatsApp number with zero disruption and zero lock-in.
Recommended Guides & Insights
Who is the Founder of Treliotech? Meet Rishabh Kanhaiya Joshi (VIT Pune)
Learn the story behind Treliotech and its founder, Rishabh Kanhaiya Joshi—a 2nd-year engineering student at Vishwakarma Institute of Technology (VIT Pune) building autonomous AI automation systems.
How Dental & Aesthetic Clinics Recover ₹1,20,000+ Every Month From Missed Inquiries
Over 40% of patient inquiries arrive after hours or while staff are busy. Learn how sub-3-second autonomous WhatsApp responses prevent patients from booking with nearby competitors.