DAEMON CATALOG 7 Modular Components • Pure Rust • Zero dynamic C libraries • Hardened systemd drop-ins

inferenced — Hardware Arbiter & Preemption Broker

Varlink Socket: /run/syntrop/io.syntrop.Inference1 • Interface: io.syntrop.Inference1 • v0.3.10

PID 1 DECOUPLED

inferenced acts as the single point of coordination for heterogeneous hardware acceleration across Linux systems. It scans DRM subsystem render nodes (/dev/dri/renderD*), NPUs, and CPU vector extensions (AMX, AVX-512) to create a dynamic unified compute plane topology.

Feature Technical Implementation Security Directives
Preemption Two-tier priority model: EmergencyTriage (100) preempts Interactive (10) and Batch (0). Freezes workloads via cgroup.freeze within 250ms. DynamicUser=yes
Zero-Copy Allocates anonymous memory fds via memfd_create(), applies write/shrink seals, passes descriptors across Unix sockets via SCM_RIGHTS. ProtectSystem=strict
Pressure Throttling Monitors /proc/pressure/{cpu,memory,io} to back off allocations before OOM killer triggers. MemoryDenyWriteExecute=yes
Gang Scheduling & Composite Leases Atomic all-or-nothing multi-accelerator allocations with PCIe tree depth and NUMA distance affinity scoring. Issues composite leases binding sequential pipeline stages or tensor parallel planes with two-tier preemption. ProtectControlGroups=no (cgroup v2 freezer)
systemd/inferenced.service
[Unit]
Description=syntropd Hardware Accelerator & Inference Arbiter
Documentation=https://syntropd.github.io/daemons.html#inferenced
Requires=inferenced.socket
After=inferenced.socket

[Service]
Type=exec
ExecStart=/usr/bin/inferenced
StandardInput=socket
Restart=on-failure
RestartSec=5s
RuntimeMaxSec=1800s

# Sandboxing
DynamicUser=yes
ProtectSystem=strict
ProtectHome=yes
PrivateTmp=yes
MemoryDenyWriteExecute=yes
RestrictAddressFamilies=AF_UNIX
DevicePolicy=closed
DeviceAllow=/dev/dri/renderD* rw

modeld — Content-Addressable Model Store

Varlink Socket: /run/syntrop/io.syntrop.Model1 • Interface: io.syntrop.Model1

CAS CACHE

modeld governs local model cache storage under /var/lib/models. All artifacts are stored by SHA-256 digest, supporting instant reflink deduplication on supported copy-on-write filesystems (Btrfs, XFS, ZFS) and hardlink trees on standard ext4 filesystems.

systemd/modeld.service
[Unit]
Description=syntropd Content-Addressable Model Store
Documentation=https://syntropd.github.io/daemons.html#modeld
Requires=modeld.socket
After=modeld.socket

[Service]
Type=exec
ExecStart=/usr/bin/modeld
StandardInput=socket
Restart=on-failure
StateDirectory=syntrop/models
ProtectSystem=strict
ProtectHome=yes
ReadOnlyPaths=/usr

contextd — System Chronology & Configuration Drift Observer

Varlink Socket: /run/syntrop/io.syntrop.Context1 • Interface: io.syntrop.Context1

CHRONOLOGY

When an outage strikes, the immediate question is always: what changed? contextd monitors system configuration directories (/etc/systemd/system, /usr/lib/systemd/system) and package manager logs to maintain a rolling microsecond-indexed ring buffer of state modifications.

toold — Sandboxed Remediation Engine & Rollback Checkpointer

Varlink Socket: /run/syntrop/io.syntrop.Tool1 • Interface: io.syntrop.Tool1

LANDLOCK SANDBOX

Autonomous self-healing requires rigorous guardrails. toold executes only strictly allowlisted, parameterized remediation actions. Each action runs in an isolated Linux Landlock sandbox with seccomp-bpf syscall filters.

Tool Name Action Mode Sandbox Restrictions Rollback Support
daemon_reload Safe (read/exec) Restricted to D-Bus signal to PID 1 N/A
unit_restart Remediation Target unit name validation via Polkit N/A
revert_dropin Remediation Landlock write limited to /etc/systemd/system/<unit>.d/ Yes (Automatic snapshot)
clear_cache Maintenance Path restricted to target transient directory No
bubblewrap_sandbox Isolation Unprivileged user/net namespaces, read-only rootfs, Landlock fallback Enforced on all tool processes

runtimed — Owned-Engine Neural Inference, Vision & Embeddings

Varlink Socket: /run/syntrop/io.syntrop.Runtime1 • Interface: io.syntrop.Runtime1 • v0.5.14

ZERO-C DEPENDENCIES (CPU BUILD)

runtimed runs its own from-scratch Rust inference engine — no llama.cpp, no cloud. It executes GGUF weights (Q8_0, Q4_K family, F16) across Gemma4 and Qwen2 architectures with differential oracle proofs, plus a vision tower for image input, live LoRA fusion, SHA256 weight pins, and normalized vector embeddings. The default CPU build links no dynamic C runtime libraries (zero libstdc++.so, zero libgomp.so); the optional cuda feature adds NVIDIA acceleration (F16-resident weights, F32-exact compute, one GPU per daemon) linking only the NVIDIA driver libraries. Fleet routers steer it via GetLoad (free slots, resident bytes).

Architecture Feature Technical Implementation Performance & Isolation
Constrained Grammar Decoding FSM token filtering supporting JSON schema, regex, and Varlink protocols with SIMD bitset masking and trie lookahead. Guaranteed syntactically valid model outputs; zero parser crashes
Heterogeneous Speculative Decoding Draft-target speculative verification engine paired with O(1) KV-cache rollback (truncate). 2x–3x throughput on mixed CPU/GPU hardware
Reasoning Token Budgeting Enforces thinking token limits via forced </think> token emission and logit masking. Strict test-time compute scaling bounds
Attention Sinks & Infinite Streaming Context StreamingLLM attention sinks (SinkWindowCache, sink_causal_mask) and bounded StreamJournal. Unbounded continuous log streaming with zero OOM risk
Intra-Node Pipeline Parallelism Multi-stage pipeline execution with direct P2P DMA activation passing across contiguous GPU ranges. Slices sealed memfd CAS layers with Gemma4 tied embeddings support. Zero-copy activation transfers; sub-millisecond stage handoffs
Hybrid GPU + Host RAM Execution Automatic fallback placing early transformer layers in GPU VRAM and overflow layers in pinned host RAM. Seamlessly integrates with Linux memory controllers. Enables execution of large parameter models exceeding single GPU VRAM
Two-Tier Paged KV Cache 16-token block paging with L1 VRAM and L2 pinned host RAM. Async spill manager monitors kernel PSI memory pressure and unspills blocks on demand. Zero memory fragmentation; preemptive PSI pressure shedding
Pure Rust Tensor Parallelism Row-parallel and column-parallel matrix sharding over NVLink using pure Rust ring all-reduce. Zero C communication runtime libraries. Linear scaling across intra-node NVLink interconnects
Real-Time Multimedia Streaming Async 24kHz S16LE PCM speech streaming via PipeWire (pw-cat) and atomic 1-step SD-Turbo generative PNG output under isolated compute leases. Sub-100ms first audio packet; atomic host runtime file writes
ldd verification (CPU build: Zero Dynamic C Runtime)
$ ldd /usr/local/bin/runtimed
        linux-vdso.so.1 (0x00007ffe341fa000)
        libgcc_s.so.1 => /lib64/libgcc_s.so.1 (0x00007f59d4d9b000)
        libm.so.6 => /lib64/libm.so.6 (0x00007f59d4c70000)
        libc.so.6 => /lib64/libc.so.6 (0x00007f59d4a8e000)
        /lib64/ld-linux-x86-64.so.2 (0x00007f59d4d9b000)
      

systemd-sentry — Autonomous Zero-Trust Supervisor

Enclave Socket: /run/syntrop/sentry.sock • Interface: io.syntrop.Sentry1

AUTONOMOUS SUPERVISOR

systemd-sentry is the proactive supervisor and brain of the autonomous self-healing loop. While inferenced, modeld, contextd, toold, and runtimed provide the passive OS capabilities and tools, sentry continuously observes systemd event loops via D-Bus, detects failures in real time, and drives the sub-200ms emergency triage pipeline.

Responsibility Technical Mechanism Failure Safeguards
Failure Interception Subscribes to org.freedesktop.systemd1.Manager.JobRemoved signals via native sd-bus event loop. Responds within 2ms. Bounded non-blocking signal buffers
Triage Pipeline Coordination Calls contextd for chronology, acquires EmergencyTriage leases from inferenced, dispatches prompt to runtimed, and executes via toold. Hard timeout budgets per phase
Circuit Breaking Enforces deterministic cooldown windows and exponential backoffs if a flapping service fails repeatedly. Never restarts units indefinitely
Multi-GPU Gang Triage Automated crash isolation for multi-accelerator gang execution failures. Discriminates DRM driver hangs, peer DMA interconnect timeouts, and VRAM exhaustion to quarantine faulty GPUs while isolating healthy devices for safe workload resumption. Targeted per-device quarantine; zero broad blast radius
systemd/systemd-sentry.service
[Unit]
Description=syntropd Autonomous Zero-Trust Supervisor
Documentation=https://syntropd.github.io/daemons.html#sentry
After=network.target dbus.socket syntrop-sockets.target
Requires=syntrop-sockets.target

[Service]
Type=notify
ExecStart=/usr/bin/systemd-sentry
Restart=always
RestartSec=3s
WatchdogSec=10s

# Hardened sandboxing
ProtectSystem=strict
ProtectHome=yes
ProtectKernelTunables=yes
MemoryDenyWriteExecute=yes
RestrictAddressFamilies=AF_UNIX AF_NETLINK

routerd — Multi-Provider LLM Router & Telemetry Reverse Proxy

Varlink Socket: /run/syntrop/io.syntrop.Router1 • Wire Ingress: 127.0.0.1:32768 & /run/syntrop/router.sock

WIRE ROUTER

routerd provides high-performance, intelligent model routing and reverse proxying across heterogeneous backends (local runtimed/inferenced, LAN runtimed peers, and external cloud LLM APIs). It incorporates real-time kernel Pressure Stall Information (PSI) and VRAM monitoring to prevent host memory starvation by seamlessly spilling compute loads to remote tiers.

Architecture Vector Technical Implementation Security & Sandboxing
Prioritization Dimensions Multi-criteria heuristic scoring factoring cost, throughput (tok/sec), context window capacity, and difficulty tiers (fast vs hard). Selects optimal provider deterministically. Slice=ai.slice
Telemetry Loop Actively samples /proc/pressure/memory and GPU VRAM leases. When memory pressure exceeds threshold (PSI > 25%), local execution is throttled and offloaded to LAN/Cloud providers. MemoryHigh=24M
MemoryMax=32M
Zero Context Contamination Strict request context isolation. Internal systemd diagnostic prompts, triage tokens, and user-space LLM traffic traverse isolated ring channels with zero shared session state. ProtectSystem=strict
NoNewPrivileges=yes
Dual-Stack Wire Ingress Listens on dual-stack TCP 127.0.0.1:32768 / [::1]:32768 and local Unix domain socket /run/syntrop/router.sock for standard OpenAI API compatibility. RestrictAddressFamilies=AF_UNIX AF_INET AF_INET6
SCMP Clustered Mesh Protocol Zero-copy binary framing over low-latency TCP (TCP_NODELAY, SO_KEEPALIVE) with TLS encryption. Tracks node health, round-trip latency, and available cluster VRAM via dynamic topology registry. Inter-node mesh isolation
Chunked Prefill & Circuit Breaker 1024-token micro-batch prefill pipelining across cluster nodes. In-flight replay buffer maintains unacknowledged chunks for atomic replay to alternate nodes upon circuit breaker trip. Zero token loss on node dropout

Operator CLI: routerctl Usage

Operators and automation scripts manage routerd via routerctl:

$ routerctl status
● routerd.service - Intelligent Model Router and Wire Protocol Gateway
     Status: OPERATIONAL (Uptime: 2h 14m)
     Wire Ingress: 127.0.0.1:32768, /run/syntrop/router.sock
     Varlink IPC:  /run/syntrop/io.syntrop.Router1
     Memory RSS:   9.12 MiB (MemoryHigh: 24M, MemoryMax: 32M)
     Kernel PSI:   Memory 0.0% some (normal)
     Providers:    3 configured, 3 healthy (runtimed-local, lan-runtimed, local-vllm)
     Total Reqs:   428 served (0 routing errors)
$ routerctl providers list
PROVIDER ID       KIND     TIER    STATUS    WEIGHT   LATENCY    MODELS
runtimed-local    runtimed fast    HEALTHY   1.0      12.4ms     qwen2.5-coder-7b
lan-runtimed      runtimed fast    HEALTHY   0.8      48.2ms     gemma4:e2b, qwen2.5:7b
local-vllm        openai   hard    HEALTHY   0.9      62.1ms     qwen3:32b
$ routerctl test runtimed-local
[TEST] Probing provider 'runtimed-local' at unix:/run/syntrop/io.syntrop.Runtime1...
[OK] Provider response received in 14.2ms.
[OK] Health check passed: HTTP 200 OK (models: 1, state: ready)
$ routerctl models
Connected Model
  qwen2.5-coder-7b
  gemma4:e2b
  gemma4:e4b

Only connected provider models are listed. Routing aliases (router:fast, router:hard, router:auto) stay valid request targets and still appear in route and /v1/models.

$ routerctl default gemma4:e4b
default → gemma4:e4b (fast + hard)
$ routerctl route --model router:fast --tokens 2048
ROUTING DECISION CANDIDATES:
  1. runtimed-local / qwen2.5-coder-7b  [SCORE: 96.4] Speed: 98.0  Cost: 100.0  Cap: 90.0  (Selected)
  2. lan-runtimed / gemma4:e2b          [SCORE: 81.2] Speed: 72.0  Cost: 100.0  Cap: 75.0
  3. local-vllm / qwen3:32b             [SCORE: 64.8] Speed: 42.0  Cost: 30.0   Cap: 99.0
systemd/routerd.service
[Unit]
Description=Intelligent Model Router and Wire Protocol Gateway
Documentation=https://github.com/syntropd/routerd
After=network.target local-fs.target
Requires=routerd.socket

[Service]
Type=notify
ExecStart=/usr/local/bin/routerd --config /etc/syntrop/routerd.toml
Restart=on-failure
RestartSec=3s
WatchdogSec=30s
Slice=ai.slice

# Memory constraints
MemoryHigh=24M
MemoryMax=32M

# Security & Sandboxing
User=syntrop
Group=syntrop
NoNewPrivileges=yes
ProtectSystem=strict
ProtectHome=yes
PrivateTmp=yes
ProtectKernelModules=yes
ProtectKernelTunables=yes
RestrictAddressFamilies=AF_UNIX AF_INET AF_INET6
RestrictNamespaces=yes
MemoryDenyWriteExecute=yes
RestrictRealtime=yes
RestrictSUIDSGID=yes
LockPersonality=yes

# cgroup v2 & systemd-oomd Protection
ManagedOOMPreference=avoid
OOMScoreAdjust=-800
OOMPolicy=stop