Subsystem Topology & Control Plane
syntropd is engineered around three fundamental tenets: unprivileged execution, asynchronous Varlink IPC over Unix domain sockets, and deterministic circuit breaking.
1. Subsystem Topology
Interactive vector diagram. Click any node below to inspect its Varlink socket, cgroup limits, and kernel primitives.
Discovers heterogeneous accelerators (DRM GPU, NPU, CPU AMX/AVX-512) and arbitrates VRAM/RAM allocations. Issues leases with priority tiers (EmergencyTriage > Interactive > Batch). Preempts batch inference workloads via cgroup.freeze during service incidents.
2. The Emergency Triage Flow (Sub-200ms Incident Recovery)
When a monitored service crashes, syntropd executes a coordinated 7-phase root-cause recovery pipeline without taking locks or blocking PID 1.
3. Zero-Idle Socket Activation Model (0 MB Footprint)
Enterprise fleets cannot afford background daemons consuming idle memory. syntropd leverages systemd socket activation so daemons exist only when actively processing requests.
Each daemon ships with paired .socket and .service units. The socket unit specifies ListenStream=/run/syntrop/io.syntrop.<Name>1 and SocketMode=0660. At system boot, systemd creates the file descriptor. No daemon process runs until a client calls connect() on the socket. When active work ceases, the daemon terminates cleanly after its idle window, returning all resident memory back to the kernel page allocator.
Port & Socket Reference
Descriptor, endpoint, and transport mappings across all syntropd socket-activated units:
| Daemon | Protocol / Transport | Endpoint / Bind Path | Socket Unit | Role & Security Directives |
|---|---|---|---|---|
| routerd | TCP (dual-stack HTTP/1.1 & HTTP/2) | 127.0.0.1:32768[::1]:32768 |
routerd.socket (FD 3 & 4) |
Wire ingress reverse proxy for OpenAI-compatible clients |
| routerd | Unix Domain Socket (Stream) | /run/syntrop/router.sock |
routerd.socket (FD 5) |
SocketMode=0666; local IPC reverse proxy endpoint |
| routerd | Varlink IPC (AF_UNIX) | /run/syntrop/io.syntrop.Router1 |
routerd.socket (FD 6) |
SocketMode=0666; dynamic scoring control plane & telemetry |
| inferenced | Varlink IPC (AF_UNIX) | /run/syntrop/io.syntrop.Inference1 |
inferenced.socket (FD 3) |
SocketMode=0666; primary hardware lease arbitration |
| inferenced | Dedicated Enclave Socket | /run/syntrop/sentry.sock |
inferenced.socket (FD 4) |
SocketMode=0660, Group=syntrop; priority triage leases |
| inferenced | Unix Domain Gateway | /run/syntrop/gateway.sock |
inferenced.socket (FD 5) |
SocketMode=0660, Group=syntrop; local OpenAI-compatible gateway endpoint |
| inferenced | SCM_RIGHTS Memfd Server | /run/syntrop/fd.sock |
inferenced.socket (FD 6) |
SocketMode=0660, Group=syntrop; zero-copy shared memory descriptors |
| modeld | Varlink IPC (AF_UNIX) | /run/syntrop/io.syntrop.Model1 |
modeld.socket (FD 3) |
SocketMode=0666; content-addressable model cache & pinning |
| contextd | Varlink IPC (AF_UNIX) | /run/syntrop/io.syntrop.Context1 |
contextd.socket (FD 3) |
SocketMode=0666; chronology & config drift tracking |
| toold | Varlink IPC (AF_UNIX) | /run/syntrop/io.syntrop.Tool1 |
toold.socket (FD 3) |
SocketMode=0666; sandboxed diagnostic & rollback execution |
| runtimed | Varlink IPC (AF_UNIX) | /run/syntrop/io.syntrop.Runtime1 |
runtimed.socket (FD 3) |
SocketMode=0666; pure-Rust neural model inference |
| systemd-sentry | Unix Domain Socket (Enclave) | /run/systemd-sentry/sentry.sock |
systemd-sentry.socket |
SocketMode=0660; crash supervisor control channel |
4. Multi-Accelerator Gang Scheduling & Clustered Mesh Protocol
The /boost subsystem upgrade introduces native multi-GPU gang scheduling, intra-node pipeline & tensor parallelism, two-tier paged KV cache memory tiers, and clustered mesh splitting.
| Subsystem | Architectural Role | Implementation Mechanism | Fault Isolation Guarantee |
|---|---|---|---|
Composite Gang Scheduling (inferenced) |
Multi-GPU arbitration & topology affinity | Atomic all-or-nothing allocation across heterogeneous accelerators. PCIe tree depth and NUMA distance scoring prioritize low-latency interconnects. Two-tier preemption freezes lower-priority gangs via cgroups v2. | Zero partial deadlock; gang rolls back completely if any single slice allocation fails. |
Pipeline & Tensor Parallelism (runtimed) |
Multi-device tensor execution & hybrid scaling | Pipelined multi-stage execution with zero-copy direct P2P DMA activation passing. Sealed memfd CAS layer slicing allows hybrid GPU + CPU RAM placement. Pure Rust NVLink tensor parallelism via row/column linear sharding and ring all-reduce. | Zero external dynamic C runtime dependencies; pure Rust math kernels with differential oracle proofs. |
Two-Tier Paged KV Cache (runtimed) |
Predictable memory budgeting & spill management | 16-token block paging allocating L1 blocks in GPU VRAM and L2 blocks in pinned host RAM. Asynchronous spill manager monitors kernel PSI memory pressure, proactively offloading blocks before OOM events. | Eliminates VRAM fragmentation; decouples sequence length from hardware memory limits. |
SCMP Clustered Mesh Protocol (routerd) |
Inter-node prompt chunking & circuit breaking | Syntrop Cluster Mesh Protocol (SCMP) zero-copy binary framing over low-latency TCP (TCP_NODELAY, SO_KEEPALIVE) with TLS encryption. 1024-token chunked prefill pipelining with circuit breakers and in-flight atomic prefill replay buffer. |
Transparent failover to replacement cluster nodes without dropping prompt tokens or context state. |
Multi-GPU Crash Isolation (systemd-sentry) |
Automated gang crash triage & quarantine | Specialized gang triage discriminating DRM driver hangs, peer DMA interconnect timeouts, and VRAM exhaustion. Automatically quarantines faulty GPUs while preserving healthy devices for uninterrupted service resumption. | Guarantees fault containment to a single accelerator without causing cascading node failures. |
5. The Seven Architectural Frontiers (Deep Systems Hardening)
Enterprise reliability, bounded latencies, and deterministic reasoning guarantees implemented natively across the daemon enclave.
| Frontier | Subsystems | Mechanisms & Primitives | Operational Guarantees |
|---|---|---|---|
| 1. Constrained / Grammar-Guided Decoding | runtimed |
Deterministic FSM token filtering supporting JSON schema, regex, and Varlink nul-terminated JSON protocols with SIMD-aligned bitset logit masking and trie lookahead. | Guaranteed syntactically valid model outputs; zero parser crashes or unparseable JSON downstream. |
| 2. Heterogeneous Speculative Decoding | inferenced & runtimed |
Dual-model speculative verification engine gang-scheduling draft models onto CPU AMX/AVX-512 engines while target models execute on discrete GPUs; O(1) PagedKvCache rollback via truncate. |
2x–3x generation throughput acceleration on mixed CPU/GPU nodes without losing accuracy. |
| 3. Autonomous Agent Loops | sentry & toold |
Closed-loop diagnostic iteration with error reflection, plan synthesis, Bubblewrap/Landlock unprivileged sandboxing, and strict circuit breaking (max 5 iterations or 30s timeout). | Autonomous failure remediation without flapping, run-away loops, or security escalation. |
| 4. Test-Time Compute & Reasoning Budgets | routerd & runtimed |
Granular reasoning token allocation (reasoning_budget, max_thinking_tokens), forced </think> token cutoff & logit masking, and streaming wire ThinkFilter separating reasoning content. |
Bounded reasoning compute per request; clean client separation of thinking chains from final answer. |
| 5. Dynamic State Compression & Infinite Streaming Context | runtimed |
StreamingLLM initial attention sinks (SinkWindowCache, sink_causal_mask), unrotated RoPE position IDs, and circular StreamJournal ring buffers with byte and entry eviction ceilings. |
Unbounded continuous journal & log stream processing with constant memory and zero OOM risk. |
| 6. End-to-End Multimedia Pipeline | routerd & runtimed |
Full bidirectional multimodal engine with vectorized spatial patch pooling (2x2, 3x3, 4x4), ephemeral vision tower lifecycle unpinning VRAM post-prefill, L2 host visual KV prefix caching, real-time 24kHz S16LE streaming audio out (Kokoro TTS) to PipeWire, and 1-step SD-Turbo generative visuals rendering. | Sub-millisecond audio streaming over PipeWire; atomic visual artifact generation; zero VRAM bloat from unpinned vision encoders. |
| 7. Dynamic Memory & Host OS Elasticity | inferenced & runtimed |
Zero-idle kernel telemetry via non-blocking fixed-stack sysfs DRM sampling (/sys/class/drm/renderD*), dual-watermark hysteresis memory governor (85% spill, 65% prefetch), recursive cgroups v2 slice priority binding (interactive preemption with 250ms deadline), and circular double-buffered JIT attention layer prefetching. |
Host OS responsiveness preserved under GPU saturation; zero unhandled VRAM allocations; sub-250ms preemption of background workloads. |