←ALL DEVELOPMENT PROJECTS
PERSONAL PRODUCT
v1.2 Active Development

(SYS.06) — PERSONAL PRODUCT — DESKTOP AUTOMATION & LOCAL AI

AutoSnap

Windows capture utility with interval screenshotting and offline Whisper speech-to-text.

YEAR: 2024–2026ROLE: Creator & Desktop Developer
AutoSnap Desktop Capture & Transcription Interface

(01) — System Overview

PROJECT
CONTEXT.

A privacy-first Windows desktop productivity application combining timed multi-source screen/browser capture, FFmpeg non-real-time video frame extraction, WASAPI system-audio loopback recording, and fully local Whisper speech-to-text transcription with mixed Taglish language support.

The Operational Bottleneck / Problem

Documenting long technical meetings, online classes, and video presentations typically forces users to send proprietary audio and screen recordings to cloud transcription services with recurring costs and privacy risks.

Engineering Solution

Created a unified Windows utility that records screens, captures internal desktop audio, and processes speech-to-text locally on-device.

(02) — Architecture & Pipeline

SYSTEM
ARCHITECTURE.

.NET 8 desktop core orchestrating NAudio WASAPI stream recording, FFmpeg CLI subprocesses, and Whisper.net inference models with local disk storage.

Data & Execution Pipeline

01. Screen / Window / WASAPI Audio Capture
→
02. NAudio Buffer / FFmpeg Process
→
03. Whisper.net Local AI Inference
→
04. Transcript Recovery Engine (30s Autosave)
→
05. Local Output (Images, TXT, SRT, VTT)

(03) — Technical Focus

ENGINEERING
HIGHLIGHTS.

HL.01

Offline Speech AI with Whisper.net

Runs speech-to-text transcription entirely on local hardware using Whisper.net, keeping sensitive meetings, lectures, and audio recordings completely private.

HL.02

WASAPI Loopback System Audio Capture

Captures desktop and browser audio streams directly through Windows WASAPI loopback with NAudio without requiring virtual audio cables or microphone bleed.

HL.03

Fast Non-Real-Time FFmpeg Video Extraction

Extracts high-resolution snapshot sequences directly from video files via FFmpeg seek operations without requiring real-time video playback.

HL.04

Taglish Language Support & Crash-Resistant Sessions

Includes specialized recognition modes for mixed Filipino-English dialogue and enforces a 30-second transcript autosave interval to protect long recordings.

(04) — Implemented Capabilities

CORE
FEATURES.

01.

Timed screenshot capture for windows, full displays, and browser tabs

02.

Native browser-tab capture using getDisplayMedia sharing permissions

03.

Background system-tray execution with minimal CPU footprint

04.

FFmpeg-powered interval snapshot and frame extraction from video files

05.

WASAPI desktop loopback audio recording with zero microphone noise

06.

Offline Whisper transcription supporting English, Tagalog, and Taglish

07.

Local Whisper Model Manager supporting Tiny, Base, Small, Medium, and Turbo

08.

Hardware-aware model recommendations matching system CPU and RAM specs

09.

Built-in transcript editor with TXT, SRT, and VTT subtitle export

10.

Periodic 30-second autosave transcript recovery for uninterrupted sessions

(05) — Technology Stack

STACK
ARCHITECTURE.

Runtime & UI

C# 12.NET 8Windows FormsSystem Tray IntegrationSelf-Contained win-x64 Build

Audio & Speech AI

Whisper.netLocal GGML ModelsNAudioWASAPI Loopback CaptureTaglish Language Mode

Video Processing

FFmpeg / FFprobe AutomationNon-Real-Time SeekingInterval Frame Extraction

Packaging & CI

Inno Setup InstallerPortable ZIP ReleasesGitHub Actions CI/CD

(06) — Problem Solving

TECHNICAL
CHALLENGES.

CHALLENGE 01

Managing Large Whisper Model Execution on Consumer PCs

Implemented a hardware-aware Whisper model recommendation engine that benchmarks available system memory and threads to prevent out-of-memory crashes.

(07) — Reliability & Governance

SECURITY & TESTING.

Privacy & Security Model

Strictly offline local media processing with zero cloud transmission.