TOPOLOGY Clean User-Space Decoupling from PID 1 • Click any architecture node below to inspect live kernel bindings

1. Subsystem Topology

Interactive vector diagram. Click any node below to inspect its Varlink socket, cgroup limits, and kernel primitives.

INTERACTIVE
LAYER 1: USER-SPACE OPERATORS & SYSTEMD SERVICES LAYER 2: SYNTROPD UNPRIVILEGED DAEMON ENCLAVE (AF_UNIX /run/syntrop/*) LAYER 3: LINUX KERNEL INTERFACES & HARDWARE ACCELERATORS syntropctl Unified Operator CLI systemd-sentry Crash Interception Enclave *.service Units Target Workloads routerd io.syntrop.Router1 127.0.0.1:32768 • Multi-Provider Proxy • Dynamic Scoring • PSI Memory Spilling ⇄ Cloud/LAN/Local inferenced io.syntrop.Inference1 gateway.sock (local) • DRM Accelerator • Two-Tier Leases • PSI Preemption • Zero-Copy Memfd modeld io.syntrop.Model1 /var/lib/models • SHA-256 CAS • Reflink Dedup • Triage Pinning • LRU Storage Quota contextd io.syntrop.Context1 /run/syntrop • Journal Slicing • Config Diffs • Event Chronology • Ring Buffer (10k) toold io.syntrop.Tool1 /run/syntrop • Landlock Sandbox • Polkit Control • Atomic Rollbacks • Seccomp Filters runtimed io.syntrop.Runtime1 In-Process Candle • In-Process GGUF • 128-dim Embed • Zero C Deps • Socket Activated Cloud & LAN Clusters LAN Peers / OpenAI / Anthropic DRM / dev-dri-* reflink / CAS fanotify / journal Landlock / seccomp cgroups v2 / memfd
inferenced.service (Compute & Arbitration Broker) ACTIVE COMPONENT

Discovers heterogeneous accelerators (DRM GPU, NPU, CPU AMX/AVX-512) and arbitrates VRAM/RAM allocations. Issues leases with priority tiers (EmergencyTriage > Interactive > Batch). Preempts batch inference workloads via cgroup.freeze during service incidents.

Varlink Socket /run/syntrop/io.syntrop.Inference1
Primary Interface io.syntrop.Inference1
Kernel Primitives
DRM/accel, cgroups v2, PSI, sealed memfd
systemd Security
DynamicUser=yes, ProtectSystem=strict

2. The Emergency Triage Flow (Sub-200ms Incident Recovery)

When a monitored service crashes, syntropd executes a coordinated 7-phase root-cause recovery pipeline without taking locks or blocking PID 1.

TOTAL BUDGET < 200ms

3. Zero-Idle Socket Activation Model (0 MB Footprint)

Enterprise fleets cannot afford background daemons consuming idle memory. syntropd leverages systemd socket activation so daemons exist only when actively processing requests.

0 MB IDLE RAM
1 Boot: .socket Only PID 1 listens on /run/syntrop/* RAM: 0 MB 2 First Call Received LISTEN_FDS handoff Startup: < 15 ms 3 Active Processing Streams Varlink replies RAM: ~12 MB 4 Inactivity Expiry Daemon calls exit(0) RAM: 0 MB Returned
HOW IT WORKS IN SYSTEMD

Each daemon ships with paired .socket and .service units. The socket unit specifies ListenStream=/run/syntrop/io.syntrop.<Name>1 and SocketMode=0660. At system boot, systemd creates the file descriptor. No daemon process runs until a client calls connect() on the socket. When active work ceases, the daemon terminates cleanly after its idle window, returning all resident memory back to the kernel page allocator.

Port & Socket Reference

Descriptor, endpoint, and transport mappings across all syntropd socket-activated units:

Daemon Protocol / Transport Endpoint / Bind Path Socket Unit Role & Security Directives
routerd TCP (dual-stack HTTP/1.1 & HTTP/2) 127.0.0.1:32768
[::1]:32768
routerd.socket (FD 3 & 4) Wire ingress reverse proxy for OpenAI-compatible clients
routerd Unix Domain Socket (Stream) /run/syntrop/router.sock routerd.socket (FD 5) SocketMode=0666; local IPC reverse proxy endpoint
routerd Varlink IPC (AF_UNIX) /run/syntrop/io.syntrop.Router1 routerd.socket (FD 6) SocketMode=0666; dynamic scoring control plane & telemetry
inferenced Varlink IPC (AF_UNIX) /run/syntrop/io.syntrop.Inference1 inferenced.socket (FD 3) SocketMode=0666; primary hardware lease arbitration
inferenced Dedicated Enclave Socket /run/syntrop/sentry.sock inferenced.socket (FD 4) SocketMode=0660, Group=syntrop; priority triage leases
inferenced Unix Domain Gateway /run/syntrop/gateway.sock inferenced.socket (FD 5) SocketMode=0660, Group=syntrop; local OpenAI-compatible gateway endpoint
inferenced SCM_RIGHTS Memfd Server /run/syntrop/fd.sock inferenced.socket (FD 6) SocketMode=0660, Group=syntrop; zero-copy shared memory descriptors
modeld Varlink IPC (AF_UNIX) /run/syntrop/io.syntrop.Model1 modeld.socket (FD 3) SocketMode=0666; content-addressable model cache & pinning
contextd Varlink IPC (AF_UNIX) /run/syntrop/io.syntrop.Context1 contextd.socket (FD 3) SocketMode=0666; chronology & config drift tracking
toold Varlink IPC (AF_UNIX) /run/syntrop/io.syntrop.Tool1 toold.socket (FD 3) SocketMode=0666; sandboxed diagnostic & rollback execution
runtimed Varlink IPC (AF_UNIX) /run/syntrop/io.syntrop.Runtime1 runtimed.socket (FD 3) SocketMode=0666; pure-Rust neural model inference
systemd-sentry Unix Domain Socket (Enclave) /run/systemd-sentry/sentry.sock systemd-sentry.socket SocketMode=0660; crash supervisor control channel

4. Multi-Accelerator Gang Scheduling & Clustered Mesh Protocol

The /boost subsystem upgrade introduces native multi-GPU gang scheduling, intra-node pipeline & tensor parallelism, two-tier paged KV cache memory tiers, and clustered mesh splitting.

DISTRIBUTED ACCELERATION
Subsystem Architectural Role Implementation Mechanism Fault Isolation Guarantee
Composite Gang Scheduling (inferenced) Multi-GPU arbitration & topology affinity Atomic all-or-nothing allocation across heterogeneous accelerators. PCIe tree depth and NUMA distance scoring prioritize low-latency interconnects. Two-tier preemption freezes lower-priority gangs via cgroups v2. Zero partial deadlock; gang rolls back completely if any single slice allocation fails.
Pipeline & Tensor Parallelism (runtimed) Multi-device tensor execution & hybrid scaling Pipelined multi-stage execution with zero-copy direct P2P DMA activation passing. Sealed memfd CAS layer slicing allows hybrid GPU + CPU RAM placement. Pure Rust NVLink tensor parallelism via row/column linear sharding and ring all-reduce. Zero external dynamic C runtime dependencies; pure Rust math kernels with differential oracle proofs.
Two-Tier Paged KV Cache (runtimed) Predictable memory budgeting & spill management 16-token block paging allocating L1 blocks in GPU VRAM and L2 blocks in pinned host RAM. Asynchronous spill manager monitors kernel PSI memory pressure, proactively offloading blocks before OOM events. Eliminates VRAM fragmentation; decouples sequence length from hardware memory limits.
SCMP Clustered Mesh Protocol (routerd) Inter-node prompt chunking & circuit breaking Syntrop Cluster Mesh Protocol (SCMP) zero-copy binary framing over low-latency TCP (TCP_NODELAY, SO_KEEPALIVE) with TLS encryption. 1024-token chunked prefill pipelining with circuit breakers and in-flight atomic prefill replay buffer. Transparent failover to replacement cluster nodes without dropping prompt tokens or context state.
Multi-GPU Crash Isolation (systemd-sentry) Automated gang crash triage & quarantine Specialized gang triage discriminating DRM driver hangs, peer DMA interconnect timeouts, and VRAM exhaustion. Automatically quarantines faulty GPUs while preserving healthy devices for uninterrupted service resumption. Guarantees fault containment to a single accelerator without causing cascading node failures.

5. The Seven Architectural Frontiers (Deep Systems Hardening)

Enterprise reliability, bounded latencies, and deterministic reasoning guarantees implemented natively across the daemon enclave.

7 FRONTIERS
Frontier Subsystems Mechanisms & Primitives Operational Guarantees
1. Constrained / Grammar-Guided Decoding runtimed Deterministic FSM token filtering supporting JSON schema, regex, and Varlink nul-terminated JSON protocols with SIMD-aligned bitset logit masking and trie lookahead. Guaranteed syntactically valid model outputs; zero parser crashes or unparseable JSON downstream.
2. Heterogeneous Speculative Decoding inferenced & runtimed Dual-model speculative verification engine gang-scheduling draft models onto CPU AMX/AVX-512 engines while target models execute on discrete GPUs; O(1) PagedKvCache rollback via truncate. 2x–3x generation throughput acceleration on mixed CPU/GPU nodes without losing accuracy.
3. Autonomous Agent Loops sentry & toold Closed-loop diagnostic iteration with error reflection, plan synthesis, Bubblewrap/Landlock unprivileged sandboxing, and strict circuit breaking (max 5 iterations or 30s timeout). Autonomous failure remediation without flapping, run-away loops, or security escalation.
4. Test-Time Compute & Reasoning Budgets routerd & runtimed Granular reasoning token allocation (reasoning_budget, max_thinking_tokens), forced </think> token cutoff & logit masking, and streaming wire ThinkFilter separating reasoning content. Bounded reasoning compute per request; clean client separation of thinking chains from final answer.
5. Dynamic State Compression & Infinite Streaming Context runtimed StreamingLLM initial attention sinks (SinkWindowCache, sink_causal_mask), unrotated RoPE position IDs, and circular StreamJournal ring buffers with byte and entry eviction ceilings. Unbounded continuous journal & log stream processing with constant memory and zero OOM risk.
6. End-to-End Multimedia Pipeline routerd & runtimed Full bidirectional multimodal engine with vectorized spatial patch pooling (2x2, 3x3, 4x4), ephemeral vision tower lifecycle unpinning VRAM post-prefill, L2 host visual KV prefix caching, real-time 24kHz S16LE streaming audio out (Kokoro TTS) to PipeWire, and 1-step SD-Turbo generative visuals rendering. Sub-millisecond audio streaming over PipeWire; atomic visual artifact generation; zero VRAM bloat from unpinned vision encoders.
7. Dynamic Memory & Host OS Elasticity inferenced & runtimed Zero-idle kernel telemetry via non-blocking fixed-stack sysfs DRM sampling (/sys/class/drm/renderD*), dual-watermark hysteresis memory governor (85% spill, 65% prefetch), recursive cgroups v2 slice priority binding (interactive preemption with 250ms deadline), and circular double-buffered JIT attention layer prefetching. Host OS responsiveness preserved under GPU saturation; zero unhandled VRAM allocations; sub-250ms preemption of background workloads.