AI & Systems·Jan 14, 2026·10 min read

Building Pyrants & MobiCam: Zero-Latency AI Surveillance & Interactive Python Engines

Yashvardhan Sharma· Founder @ Pyrants & Gati Music

Executive Overview

MobiCam (GitHub Repository) and Pyrants address two core engineering challenges:

  1. MobiCam: Converting any Android smartphone or USB camera into a high-performance, zero-latency AI security camera. It pairs a native open-source Android app (MyCam Server built with Kotlin & CameraX) with a Python AI client using OpenCV, MobileNet-SSD deep neural networks, and a continuous socket-flushing buffer (FastStreamReader).
  2. Pyrants: An interactive Python learning dashboard built with Next.js 16 that runs sandboxed code evaluation, streams real-time execution outputs, and provides computer vision exercise labs.

By combining low-level socket stream control with on-device deep learning object detection, MobiCam eliminates stream delay while supporting remote global monitoring over 4G/5G via Tailscale VPN.

System Architecture & Workflow Charts

Here is the exact data flow for MobiCam's zero-latency video processing and Pyrants' interactive code execution engine.

MobiCam Architecture Pipeline

text
+-----------------------------------------------------------------------------------+
|                        MOBILE CAMERA NODE (ANDROID / IOS)                         |
|  [ Android Phone Camera ]                                                         |
|          | (CameraX Frame Ingest @ 720p/480p 30 FPS)                             |
|          v                                                                        |
|  [ MyCam Server App ] (MjpegStreamServer.kt on HTTP:8080)                        |
+----------|------------------------------------------------------------------------+
           | (MJPEG Stream over Wi-Fi / USB / Tailscale 4G/5G Mesh)
           v
+-----------------------------------------------------------------------------------+
|                         DESKTOP AI CLIENT (PYTHON / OPENCV)                       |
|  [ FastStreamReader Thread ]                                                      |
|          |---> Flushes Socket Buffer Continuously (0-Lag Frame Pull)               |
|          v                                                                        |
|  [ AI Engine: MobileNet-SSD Caffe DNN & HOG Detector ]                            |
|          |---> Classifies Objects: Person, Pets, Vehicles                         |
|          |---> Calculates Bounding Box Overlays                                   |
|          v                                                                        |
|  [ Event Triggers & Actions ]                                                     |
|          +---> Auto 5-Second Video Clips (.mp4) -> recordings/                    |
|          +---> Auto High-Res Person Snapshots (.jpg) -> snapshots/                |
|          +---> Audible Security Audio Alarm Trigger                               |
|          v                                                                        |
|  [ Main GUI Viewport & CLI Window ] (main.py / cli_cam.py)                        |
+-----------------------------------------------------------------------------------+

Pyrants Interactive Sandbox Pipeline

text
+-----------------------------------------------------------------------------------+
|                           NEXT.JS 16 DASHBOARD (FRONTEND)                         |
|  [ Interactive Student IDE Editor ]                                               |
|          | (User Python Code Payload)                                             |
|          v                                                                        |
|  [ WebSockets Gateway Client ]                                                    |
+----------|------------------------------------------------------------------------+
           | (JSON Payload over WSS)
           v
+-----------------------------------------------------------------------------------+
|                        PYRANTS EVALUATION ENGINE (BACKEND)                        |
|  [ Python AST Security Validator ] ---> Blocks Restricted AST Nodes               |
|          | (Clean Code Tree)                                                      |
|          v                                                                        |
|  [ Sandboxed Worker Scope ] ---> Redirects Stdout / Stderr Buffers               |
|          |---> Enforces 3s Execution Timeout Guard                                |
|          v                                                                        |
|  [ Live Telemetry Stream ] ---> Pushes Terminal Outputs to Frontend              |
+-----------------------------------------------------------------------------------+

How MobiCam Works: Under the Hood

1. Zero-Latency Socket Flushing (FastStreamReader)

Standard OpenCV cv2.VideoCapture or HTTP requests buffer incoming network frames in an internal queue. Over Wi-Fi or cellular networks, this buffer creates an escalating 2-5 second video delay.

MobiCam solves this with a custom threading class called FastStreamReader:

  • Runs an asynchronous daemon thread that constantly reads from the HTTP MJPEG stream socket.
  • Overwrites the latest frame in memory and discards older backlog frames instantly.
  • When the AI engine requests a frame, it always receives the instantaneous, 0-latency frame.

2. MobileNet-SSD Deep Neural Network (ai_engine.py)

  • Utilizes Caffe framework pretrained weights (MobileNetSSD_deploy.caffemodel) combined with OpenCV's cv2.dnn.readNetFromCaffe.
  • Filters detections into configurable target classes: Person Only, Pets (dogs, cats), and Vehicles (cars, buses, motorbikes).
  • Draws real-time bounding boxes, class labels, and confidence percentage overlays.

3. Automated Person Detection & Recording Triggers

  • Auto 5-Second Clips: When a person enters the camera field of view, MobiCam automatically initializes an OpenCV VideoWriter to capture a 5-second .mp4 clip to recordings/.
  • Auto Snapshots & Alarms: Saves high-resolution .jpg images to snapshots/ and events/, and triggers audible sound alerts for perimeter security.

4. Native Android App (MyCam Server)

  • Built using Kotlin, Android CameraX, and Jetpack Compose.
  • Hosts an embedded HTTP MJPEG stream server (MjpegStreamServer.kt) on port 8080.
  • Includes Screen Dimmer feature to prevent screen burn-in and phone overheating during 24/7 continuous operation.

Real-World Uses & Setup Guidelines

  • 24/7 Smart Home Surveillance: Turn spare Android smartphones into active home security cameras. Setting battery usage to *Unrestricted*, configuring static IPs, and dimming screens allows non-stop monitoring.
  • Global Remote Monitoring via Tailscale: Stream live mobile camera feeds anywhere in the world over 4G/5G cellular data using a Tailscale encrypted mesh network.
  • Privacy-First Local AI: All object detection and video encoding run locally on your desktop or edge device. No unencrypted video is transmitted to third-party cloud servers.
  • Interactive CV Pedagogy on Pyrants: Students can inspect live detection vectors, modify confidence thresholds, and write custom Python frame processing scripts.

Authentic Source Code Snippets

Snippet 1: Zero-Latency Socket Reader (FastStreamReader in Python)

python
import cv2
import threading
import urllib.request
import numpy as np

class FastStreamReader:
    """Continuously flushes socket frame buffers to achieve 0-latency streaming."""
    def __init__(self, url: str):
        self.url = url
        self.stream = urllib.request.urlopen(url)
        self.bytes = b''
        self.frame = None
        self.stopped = False
        self.lock = threading.Lock()

        # Start background frame reading thread
        self.thread = threading.Thread(target=self._update, daemon=True)
        self.thread.start()

    def _update(self):
        while not self.stopped:
            try:
                self.bytes += self.stream.read(4096)
                a = self.bytes.find(b'ÿØ') # JPEG start marker
                b = self.bytes.find(b'ÿÙ') # JPEG end marker

                if a != -1 and b != -1:
                    jpg = self.bytes[a:b+2]
                    self.bytes = self.bytes[b+2:] # Flush consumed buffer
                    img = cv2.imdecode(np.frombuffer(jpg, dtype=np.uint8), cv2.IMREAD_COLOR)

                    with self.lock:
                        self.frame = img
            except Exception as e:
                print(f"[FastStreamReader Error]: {e}")
                break

    def read(self):
        with self.lock:
            return self.frame is not None, self.frame

    def stop(self):
        self.stopped = True

Snippet 2: MobileNet-SSD Object Detection Engine (ai_engine.py)

python
import cv2
import numpy as np
import time

CLASSES = ["background", "aeroplane", "bicycle", "bird", "boat",
           "bottle", "bus", "car", "cat", "chair", "cow", "diningtable",
           "dog", "horse", "motorbike", "person", "pottedplant", "sheep",
           "sofa", "train", "tvmonitor"]

class MobiCamAIEngine:
    def __init__(self, prototxt_path: str, model_path: str, confidence_thresh=0.5):
        self.net = cv2.dnn.readNetFromCaffe(prototxt_path, model_path)
        self.confidence_thresh = confidence_thresh

    def detect_objects(self, frame: np.ndarray):
        (h, w) = frame.shape[:2]
        # Prepare 300x300 blob for MobileNet-SSD DNN
        blob = cv2.dnn.blobFromImage(cv2.resize(frame, (300, 300)), 0.007843, (300, 300), 127.5)
        self.net.setInput(blob)
        detections = self.net.forward()

        results = []
        for i in range(detections.shape[2]):
            confidence = detections[0, 0, i, 2]
            if confidence > self.confidence_thresh:
                idx = int(detections[0, 0, i, 1])
                label = CLASSES[idx]
                box = detections[0, 0, i, 3:7] * np.array([w, h, w, h])
                (startX, startY, endX, endY) = box.astype("int")

                results.append({
                    "label": label,
                    "confidence": float(confidence),
                    "box": (startX, startY, endX, endY)
                })

                # Draw visual bounding box and label overlay
                cv2.rectangle(frame, (startX, startY), (endX, endY), (0, 255, 0), 2)
                text = f"{label}: {confidence * 100:.1f}%"
                cv2.putText(frame, text, (startX, startY - 10),
                            cv2.FONT_HERSHEY_SIMPLEX, 0.5, (0, 255, 0), 2)

        return results, frame

Snippet 3: Pyrants AST Security & Code Execution Sandbox (Python)

python
import ast
import sys
import io

FORBIDDEN_NODES = {ast.Import, ast.ImportFrom, ast.Global, ast.Nonlocal}

def validate_python_ast(code_str: str) -> bool:
    """Validates student Python submission for restricted AST calls."""
    try:
        tree = ast.parse(code_str)
        for node in ast.walk(tree):
            if type(node) in FORBIDDEN_NODES:
                return False
            if isinstance(node, ast.Call) and getattr(node.func, 'id', '') in {'exec', 'eval', 'open', '__import__'}:
                return False
        return True
    except SyntaxError:
        return False

def run_pyrants_sandbox(code_str: str) -> dict:
    if not validate_python_ast(code_str):
        return {"status": "error", "message": "Restricted call detected by Pyrants AST validator."}

    buffer = io.StringIO()
    sys.stdout = buffer
    sys.stderr = buffer
    safe_scope = {"__builtins__": {"print": print, "range": range, "len": len, "int": int, "str": str}}

    try:
        exec(code_str, safe_scope)
        return {"status": "success", "output": buffer.getvalue()}
    except Exception as err:
        return {"status": "runtime_error", "output": str(err)}
    finally:
        sys.stdout = sys.__stdout__
        sys.stderr = sys.__stderr__

Project Credits & Official Resources

  • Creator & Lead Architect: Yashvardhan Sharma (@Yashvardhan4646)
  • MobiCam Repository: GitHub - Yashvardhan4646/MobiCam
  • Official Native Android App: MyCam Server (MyCamServer.apk in releases/), built with Kotlin, Android CameraX, and Jetpack Compose.
  • Core Python AI Stack: Python 3.8+, OpenCV 4.5+, MobileNet-SSD (Caffe DNN framework), NumPy.
  • Web & Dashboard Infrastructure: Next.js 16 (App Router), React 19, TypeScript, Tailwind CSS, Supabase.
  • License: MIT License (Free for personal and commercial open-source use).
#Python#OpenCV#MobileNet-SSD#Android#Kotlin#AI Surveillance#Next.js 16