MAP Runtime: Modular Attention and Perception Runtime by Randy JohnsonMAP Runtime: Modular Attention and Perception Runtime by Randy Johnson

MAP Runtime: Modular Attention and Perception Runtime

Randy Johnson

Randy Johnson

MAP Runtime: Modular Attention and Perception Runtime
The problem. AI agents can reason about a prompt, but they do not naturally share a reliable understanding of what is happening around them. Cameras, microphones, screens, terminals, serial devices, and USB hardware all produce asynchronous, noisy, short-lived signals. Without a common clock, retained evidence, source health, and explicit uncertainty, an agent cannot reliably determine what happened, when it happened, whether two observations belong together, or why a conclusion should be trusted. MAP provides that missing perception and evidence layer between the live world and downstream intelligence.
The core layer. MAP converts physical and digital activity into a synchronized, inspectable model of observable reality. It captures multimodal streams, maintains bounded sensory buffers, performs local perception, directs compute toward meaningful changes, correlates events across sources, preserves evidence around important moments, and records the resulting timeline for search and replay. Every event retains its timing, source, confidence, provenance, and connection to supporting media. Agents and reasoning systems receive grounded context rather than an undifferentiated stream, and they can request nearby evidence when the initial context is insufficient. Sources, perception models, storage, memory, reasoning, and speech providers remain replaceable behind narrow contracts, while personality, beliefs, and application-specific meaning stay outside MAP.
The implementation. MAP is a completed local-first Python 3.12 runtime with a frozen plugin API, typed provider SDK, health reporting, permission enforcement, retention policies, recovery handling, SQLite event history, filesystem media storage, and an interactive replay workbench. Its verified providers include OpenCV camera perception, microphone capture, CPU Silero voice activity detection, GPU faster-whisper transcription, screen capture, terminal observation, serial monitoring, and USB device-state tracking. Optional adapters connect preserved evidence to KGK knowledge promotion, MAR reasoning, and Chatterbox Turbo speech output without making any of them core dependencies. The full ReSpeaker/XIAO reference workflow has been exercised against real hardware—including physical disconnect, reconnect, recovery, evidence preservation, reasoning expansion, and replay—without requiring Docker or cloud services.
Like this project

Posted Sep 18, 2026

A local-first perception runtime that turns live physical and digital signals into synchronized, inspectable evidence for AI agents.