The UNIX Way Meets Modern Machine Intelligence
AI belongs in user space, cleanly decoupled from PID 1. Rather than bloating the init system, syntropd adheres to the systemd modular architecture: small, isolated daemons with clearly bounded responsibilities communicating over standard Varlink sockets.
Heterogeneous accelerator discovery and compute lease arbiter. Discovers DRM GPU, NPU, and CPU AMX/AVX-512 extensions. Governs memory pressure via cgroups v2 and preempts batch workloads during system emergencies.
- Dynamic DRM render node discovery
- Two-tier preemption (EmergencyTriage > Batch)
- Kernel PSI pressure monitoring (CPU, I/O, VRAM)
- Sealed memfd zero-copy tensor passing
Content-Addressable Model Store under /var/lib/models. Deduplicates weights via filesystem reflinks, validates SHA-256 digests, and guarantees emergency triage models remain pinned in local storage.
- SHA-256 CAS repository with tag aliases
- Reflink deduplication on Btrfs/XFS/ZFS
- Autonomous LRU cache quota management
- Zero-eviction pinning for triage models
Causality observer and configuration drift tracker. Observes unit drop-ins, package transactions (dnf/apt/pacman), and systemd journal events to correlate the exact chain of events leading to a failure.
- fanotify / inotify tracking of
/etc/systemd/system - Microsecond-precision causality ring buffer
- Unified diff snapshotting before/after changes
- Structured incident timeline extraction
Sandboxed execution engine for safe diagnostic and remediating actions. Restricts tool execution using Linux Landlock LSM and seccomp filters, creating atomic filesystem snapshots before every mutation.
- Strict execution allowlist & Polkit authentication
- Landlock LSM unprivileged sandboxing
- Automatic pre-execution rollback checkpoints
- One-step deterministic remediation rollback
Owned-engine neural inference: text, vision, and embeddings with zero dynamic C dependencies in the CPU build, lazy socket activation, and zero background idle memory.
- In-VRAM 4-bit Marlin GEMV & INT4 tiling
- Ada Lovelace FP8 Tensor Core pipeline
- Quantized paged KV-cache (16-token FP8 blocks)
- Dual-GPU speculative decoding offload
- GGUF engine: Gemma4/Qwen2, vision, LoRA
- Socket-activated lifecycle with clean exit
Multi-provider LLM reverse proxy and dynamic scoring router. Distributes inference across local runtimed/inferenced instances, LAN runtimed peers, and cloud providers with kernel PSI offloading and zero context contamination.
- Dynamic candidate scoring (cost, throughput, context, tiers)
- Kernel PSI telemetry loop & VRAM pressure offloading
- Dual-stack TCP ingress (port 32768) & /run/syntrop/router.sock
- Native Varlink IPC control socket (io.syntrop.Router1)
Unified companion CLI tool. Queries daemon health, introspects hardware accelerators, inspects configuration drift diffs, generates shell completions, and formats diagnostic outputs cleanly for humans and JSON consumers.
syntropctl status- Health & latency probesyntropctl explain <unit>- Automated root-causesyntropctl devices- DRM GPU / NPU topologiessyntropctl drift- Config chronology diffs
The Seven Architectural Frontiers
Hardening autonomous operating system intelligence with zero compromise on safety, throughput, or memory bounds.
Deterministic FSM token filtering supporting JSON schema, regex, and Varlink nul-terminated JSON protocols with SIMD-aligned bitset logit masking and vocab trie pruning.
- Guaranteed syntactically valid JSON & Varlink payloads
- SIMD bitset token masking before Softmax
- Trie-based prefix lookahead over vocabularies
Heterogeneous speculative decoding gang-scheduling draft models onto CPU AMX/AVX-512 engines while target models run on discrete GPU/NPU planes, with O(1) KV-cache truncation.
- PlaneRole::Draft (CPU) & PlaneRole::Target (GPU)
- O(1) PagedKvCache rollback without re-computation
- 2x–3x generation throughput on mixed hardware
Closed-loop diagnostic agent iterations with error reflection, circuit breaking (max 5 iterations or 30s timeout), and Bubblewrap/Landlock unprivileged sandbox isolation.
- Bubblewrap unprivileged namespaces isolation
- Autonomous plan generation & tool execution
- Deterministic circuit-breaker halting flapping
Test-time compute scaling with reasoning budgets (reasoning_budget, max_thinking_tokens), forced </think> token cutoff, and streaming wire think separation.
- Granular reasoning token allocation per request
- ThinkFilter streaming separation of reasoning_content
- Forced </think> token injection & logit masking
Dynamic state compression with pinned attention sinks (SinkWindowCache, sink_causal_mask) and unrotated RoPE for infinite context generation, paired with bounded StreamJournal to prevent OOM in 24/7 continuous log monitoring.
- Pinned initial sink tokens preventing perplexity collapse
- Rolling FIFO context window sliding without memory growth
- Bounded StreamJournal circular ring buffer with strict byte ceilings
Full bidirectional multimodal engine: vectorized spatial patch pooling (2x2, 3x3, 4x4), ephemeral vision tower unpinning VRAM post-prefill, real-time 24kHz S16LE audio streaming to PipeWire, and 1-step SD-Turbo generative rendering.
- Vectorized spatial patch pooling & ephemeral vision unpinning
- Real-time 24kHz S16LE streaming audio out to PipeWire (pw-cat)
- 1-step SD-Turbo visual generation via isolated compute leases
Zero-idle kernel telemetry via non-blocking fixed-stack sysfs DRM sampling (/sys/class/drm/renderD*), dual-watermark hysteresis memory governor (spill from 85% high down to 70%, prefetch below 65%), recursive cgroups v2 slice priority binding, and circular double-buffered JIT attention layer prefetching.
- Non-blocking DRM sysfs telemetry (/sys/class/drm/renderD*)
- Dual-watermark hysteresis memory governor (85% spill, 65% prefetch)
- Recursive cgroups v2 interactive preemption & circular JIT prefetch
Inspect & Diagnose with syntropctl
Deterministic command-line tools modeled directly after systemctl and journalctl.
DAEMON SOCKET STATE LATENCY INTERFACES
inferenced /run/syntrop/io.syntrop.Inference1 ACTIVE 0.41ms io.syntrop.Inference1, org.varlink.service
modeld /run/syntrop/io.syntrop.Model1 ACTIVE 0.32ms io.syntrop.Model1, org.varlink.service
contextd /run/syntrop/io.syntrop.Context1 ACTIVE 0.28ms io.syntrop.Context1, org.varlink.service
toold /run/syntrop/io.syntrop.Tool1 ACTIVE 0.35ms io.syntrop.Tool1, org.varlink.service
runtimed /run/syntrop/io.syntrop.Runtime1 STANDBY 0.15ms io.syntrop.Runtime1, org.varlink.service
routerd /run/syntrop/io.syntrop.Router1 ACTIVE 0.38ms io.syntrop.Router1, org.varlink.service
System Status: OPERATIONAL (All sockets listening; 0 MB idle RAM footprint)
● payment-service.service - Core Payment Processing Daemon
Loaded: loaded (/etc/systemd/system/payment-service.service; enabled)
Active: failed (Result: exit-code) since Wed 2026-09-24 05:14:22 UTC; 18s ago
Main PID: 84920 (code=exited, status=1/FAILURE)
TRIAGE ANALYSIS (Root-Cause Correlation):
├─ Chronology: /etc/systemd/system/payment-service.service.d/override.conf modified 2m ago
├─ Config Drift: MemoryMax was reduced from 8G -> 512M
├─ Kernel PSI: Memory pressure spiked to 92.4% prior to SIGKILL
└─ Conclusion: Out-of-memory termination caused by insufficient cgroup limit.
RECOMMENDED REMEDIATION:
Execute safe drop-in rollback checkpoint:
$ syntropctl run rollback rb-84920-1a
Or restore limit:
# systemctl revert payment-service.service