inferenced acts as the single point of coordination for heterogeneous hardware acceleration across Linux systems. It scans DRM subsystem render nodes (/dev/dri/renderD*), NPUs, and CPU vector extensions (AMX, AVX-512) to create a dynamic unified compute plane topology.
| Feature |
Technical Implementation |
Security Directives |
| Preemption |
Two-tier priority model: EmergencyTriage (100) preempts Interactive (10) and Batch (0). Freezes workloads via cgroup.freeze within 250ms. |
DynamicUser=yes |
| Zero-Copy |
Allocates anonymous memory fds via memfd_create(), applies write/shrink seals, passes descriptors across Unix sockets via SCM_RIGHTS. |
ProtectSystem=strict |
| Pressure Throttling |
Monitors /proc/pressure/{cpu,memory,io} to back off allocations before OOM killer triggers. |
MemoryDenyWriteExecute=yes |
| Gang Scheduling & Composite Leases |
Atomic all-or-nothing multi-accelerator allocations with PCIe tree depth and NUMA distance affinity scoring. Issues composite leases binding sequential pipeline stages or tensor parallel planes with two-tier preemption. |
ProtectControlGroups=no (cgroup v2 freezer) |
[Unit]
Description=syntropd Hardware Accelerator & Inference Arbiter
Documentation=https://syntropd.github.io/daemons.html#inferenced
Requires=inferenced.socket
After=inferenced.socket
[Service]
Type=exec
ExecStart=/usr/bin/inferenced
StandardInput=socket
Restart=on-failure
RestartSec=5s
RuntimeMaxSec=1800s
# Sandboxing
DynamicUser=yes
ProtectSystem=strict
ProtectHome=yes
PrivateTmp=yes
MemoryDenyWriteExecute=yes
RestrictAddressFamilies=AF_UNIX
DevicePolicy=closed
DeviceAllow=/dev/dri/renderD* rw
modeld governs local model cache storage under /var/lib/models. All artifacts are stored by SHA-256 digest, supporting instant reflink deduplication on supported copy-on-write filesystems (Btrfs, XFS, ZFS) and hardlink trees on standard ext4 filesystems.
- Atomic Staging: Writes incoming models to
O_TMPFILE anonymous files and links them atomically only after full SHA-256 digest verification.
- Remote Pull & Register:
syn pull <model> streams GGUF models directly from Hugging Face or registry, calculates SHA-256 in-flight, registers via io.syntrop.Model1.Register, and triggers io.syntrop.Router1.Reload.
- Triage Pinning: Emergency reasoning models (e.g.
qwen2.5-coder-1.5b:pinned) are marked immune to LRU cache pruning.
- Segment Validation: Rejects path traversal identifiers (
., .., embedded slashes) at the Varlink boundary with InvalidIdentifier.
[Unit]
Description=syntropd Content-Addressable Model Store
Documentation=https://syntropd.github.io/daemons.html#modeld
Requires=modeld.socket
After=modeld.socket
[Service]
Type=exec
ExecStart=/usr/bin/modeld
StandardInput=socket
Restart=on-failure
StateDirectory=syntrop/models
ProtectSystem=strict
ProtectHome=yes
ReadOnlyPaths=/usr
When an outage strikes, the immediate question is always: what changed? contextd monitors system configuration directories (/etc/systemd/system, /usr/lib/systemd/system) and package manager logs to maintain a rolling microsecond-indexed ring buffer of state modifications.
- Filesystem Monitoring: Uses
fanotify and inotify to capture before/after unified diffs when unit drop-ins are edited.
- Journal Cursor Extraction: Extracts precise micro-slices of systemd logs surrounding a unit failure without expensive full-journal scans.
- Correlation Engine: Merges package updates, configuration diffs, and kernel events into a unified causality timeline returned via
GetUnitContext.
Autonomous self-healing requires rigorous guardrails. toold executes only strictly allowlisted, parameterized remediation actions. Each action runs in an isolated Linux Landlock sandbox with seccomp-bpf syscall filters.
| Tool Name |
Action Mode |
Sandbox Restrictions |
Rollback Support |
daemon_reload |
Safe (read/exec) |
Restricted to D-Bus signal to PID 1 |
N/A |
unit_restart |
Remediation |
Target unit name validation via Polkit |
N/A |
revert_dropin |
Remediation |
Landlock write limited to /etc/systemd/system/<unit>.d/ |
Yes (Automatic snapshot) |
clear_cache |
Maintenance |
Path restricted to target transient directory |
No |
bubblewrap_sandbox |
Isolation |
Unprivileged user/net namespaces, read-only rootfs, Landlock fallback |
Enforced on all tool processes |
runtimed runs its own from-scratch Rust inference engine — no llama.cpp, no cloud. It executes GGUF weights (Q8_0, Q4_K family, F16) across Gemma4 and Qwen2 architectures with differential oracle proofs, plus a vision tower for image input, live LoRA fusion, SHA256 weight pins, and normalized vector embeddings. The default CPU build links no dynamic C runtime libraries (zero libstdc++.so, zero libgomp.so); the optional cuda feature adds NVIDIA acceleration (F16-resident weights, F32-exact compute, one GPU per daemon) linking only the NVIDIA driver libraries. Fleet routers steer it via GetLoad (free slots, resident bytes).
| Architecture Feature |
Technical Implementation |
Performance & Isolation |
| Constrained Grammar Decoding |
FSM token filtering supporting JSON schema, regex, and Varlink protocols with SIMD bitset masking and trie lookahead. |
Guaranteed syntactically valid model outputs; zero parser crashes |
| Heterogeneous Speculative Decoding |
Draft-target speculative verification engine paired with O(1) KV-cache rollback (truncate). |
2x–3x throughput on mixed CPU/GPU hardware |
| Reasoning Token Budgeting |
Enforces thinking token limits via forced </think> token emission and logit masking. |
Strict test-time compute scaling bounds |
| Attention Sinks & Infinite Streaming Context |
StreamingLLM attention sinks (SinkWindowCache, sink_causal_mask) and bounded StreamJournal. |
Unbounded continuous log streaming with zero OOM risk |
| Intra-Node Pipeline Parallelism |
Multi-stage pipeline execution with direct P2P DMA activation passing across contiguous GPU ranges. Slices sealed memfd CAS layers with Gemma4 tied embeddings support. |
Zero-copy activation transfers; sub-millisecond stage handoffs |
| Hybrid GPU + Host RAM Execution |
Automatic fallback placing early transformer layers in GPU VRAM and overflow layers in pinned host RAM. Seamlessly integrates with Linux memory controllers. |
Enables execution of large parameter models exceeding single GPU VRAM |
| Two-Tier Paged KV Cache |
16-token block paging with L1 VRAM and L2 pinned host RAM. Async spill manager monitors kernel PSI memory pressure and unspills blocks on demand. |
Zero memory fragmentation; preemptive PSI pressure shedding |
| Pure Rust Tensor Parallelism |
Row-parallel and column-parallel matrix sharding over NVLink using pure Rust ring all-reduce. Zero C communication runtime libraries. |
Linear scaling across intra-node NVLink interconnects |
| Real-Time Multimedia Streaming |
Async 24kHz S16LE PCM speech streaming via PipeWire (pw-cat) and atomic 1-step SD-Turbo generative PNG output under isolated compute leases. |
Sub-100ms first audio packet; atomic host runtime file writes |
$ ldd /usr/local/bin/runtimed
linux-vdso.so.1 (0x00007ffe341fa000)
libgcc_s.so.1 => /lib64/libgcc_s.so.1 (0x00007f59d4d9b000)
libm.so.6 => /lib64/libm.so.6 (0x00007f59d4c70000)
libc.so.6 => /lib64/libc.so.6 (0x00007f59d4a8e000)
/lib64/ld-linux-x86-64.so.2 (0x00007f59d4d9b000)
systemd-sentry is the proactive supervisor and brain of the autonomous self-healing loop. While inferenced, modeld, contextd, toold, and runtimed provide the passive OS capabilities and tools, sentry continuously observes systemd event loops via D-Bus, detects failures in real time, and drives the sub-200ms emergency triage pipeline.
| Responsibility |
Technical Mechanism |
Failure Safeguards |
| Failure Interception |
Subscribes to org.freedesktop.systemd1.Manager.JobRemoved signals via native sd-bus event loop. Responds within 2ms. |
Bounded non-blocking signal buffers |
| Triage Pipeline Coordination |
Calls contextd for chronology, acquires EmergencyTriage leases from inferenced, dispatches prompt to runtimed, and executes via toold. |
Hard timeout budgets per phase |
| Circuit Breaking |
Enforces deterministic cooldown windows and exponential backoffs if a flapping service fails repeatedly. |
Never restarts units indefinitely |
| Multi-GPU Gang Triage |
Automated crash isolation for multi-accelerator gang execution failures. Discriminates DRM driver hangs, peer DMA interconnect timeouts, and VRAM exhaustion to quarantine faulty GPUs while isolating healthy devices for safe workload resumption. |
Targeted per-device quarantine; zero broad blast radius |
[Unit]
Description=syntropd Autonomous Zero-Trust Supervisor
Documentation=https://syntropd.github.io/daemons.html#sentry
After=network.target dbus.socket syntrop-sockets.target
Requires=syntrop-sockets.target
[Service]
Type=notify
ExecStart=/usr/bin/systemd-sentry
Restart=always
RestartSec=3s
WatchdogSec=10s
# Hardened sandboxing
ProtectSystem=strict
ProtectHome=yes
ProtectKernelTunables=yes
MemoryDenyWriteExecute=yes
RestrictAddressFamilies=AF_UNIX AF_NETLINK
routerd provides high-performance, intelligent model routing and reverse proxying across heterogeneous backends (local runtimed/inferenced, LAN runtimed peers, and external cloud LLM APIs). It incorporates real-time kernel Pressure Stall Information (PSI) and VRAM monitoring to prevent host memory starvation by seamlessly spilling compute loads to remote tiers.
| Architecture Vector |
Technical Implementation |
Security & Sandboxing |
| Prioritization Dimensions |
Multi-criteria heuristic scoring factoring cost, throughput (tok/sec), context window capacity, and difficulty tiers (fast vs hard). Selects optimal provider deterministically. |
Slice=ai.slice |
| Telemetry Loop |
Actively samples /proc/pressure/memory and GPU VRAM leases. When memory pressure exceeds threshold (PSI > 25%), local execution is throttled and offloaded to LAN/Cloud providers. |
MemoryHigh=24M
MemoryMax=32M |
| Zero Context Contamination |
Strict request context isolation. Internal systemd diagnostic prompts, triage tokens, and user-space LLM traffic traverse isolated ring channels with zero shared session state. |
ProtectSystem=strict
NoNewPrivileges=yes |
| Dual-Stack Wire Ingress |
Listens on dual-stack TCP 127.0.0.1:32768 / [::1]:32768 and local Unix domain socket /run/syntrop/router.sock for standard OpenAI API compatibility. |
RestrictAddressFamilies=AF_UNIX AF_INET AF_INET6 |
| SCMP Clustered Mesh Protocol |
Zero-copy binary framing over low-latency TCP (TCP_NODELAY, SO_KEEPALIVE) with TLS encryption. Tracks node health, round-trip latency, and available cluster VRAM via dynamic topology registry. |
Inter-node mesh isolation |
| Chunked Prefill & Circuit Breaker |
1024-token micro-batch prefill pipelining across cluster nodes. In-flight replay buffer maintains unacknowledged chunks for atomic replay to alternate nodes upon circuit breaker trip. |
Zero token loss on node dropout |
Operator CLI: routerctl Usage
Operators and automation scripts manage routerd via routerctl:
● routerd.service - Intelligent Model Router and Wire Protocol Gateway
Status: OPERATIONAL (Uptime: 2h 14m)
Wire Ingress: 127.0.0.1:32768, /run/syntrop/router.sock
Varlink IPC: /run/syntrop/io.syntrop.Router1
Memory RSS: 9.12 MiB (MemoryHigh: 24M, MemoryMax: 32M)
Kernel PSI: Memory 0.0% some (normal)
Providers: 3 configured, 3 healthy (runtimed-local, lan-runtimed, local-vllm)
Total Reqs: 428 served (0 routing errors)
PROVIDER ID KIND TIER STATUS WEIGHT LATENCY MODELS
runtimed-local runtimed fast HEALTHY 1.0 12.4ms qwen2.5-coder-7b
lan-runtimed runtimed fast HEALTHY 0.8 48.2ms gemma4:e2b, qwen2.5:7b
local-vllm openai hard HEALTHY 0.9 62.1ms qwen3:32b
[TEST] Probing provider 'runtimed-local' at unix:/run/syntrop/io.syntrop.Runtime1...
[OK] Provider response received in 14.2ms.
[OK] Health check passed: HTTP 200 OK (models: 1, state: ready)
Connected Model
qwen2.5-coder-7b
gemma4:e2b
gemma4:e4b
Only connected provider models are listed. Routing aliases
(router:fast, router:hard, router:auto)
stay valid request targets and still appear in route and /v1/models.
default → gemma4:e4b (fast + hard)
ROUTING DECISION CANDIDATES:
1. runtimed-local / qwen2.5-coder-7b [SCORE: 96.4] Speed: 98.0 Cost: 100.0 Cap: 90.0 (Selected)
2. lan-runtimed / gemma4:e2b [SCORE: 81.2] Speed: 72.0 Cost: 100.0 Cap: 75.0
3. local-vllm / qwen3:32b [SCORE: 64.8] Speed: 42.0 Cost: 30.0 Cap: 99.0
[Unit]
Description=Intelligent Model Router and Wire Protocol Gateway
Documentation=https://github.com/syntropd/routerd
After=network.target local-fs.target
Requires=routerd.socket
[Service]
Type=notify
ExecStart=/usr/local/bin/routerd --config /etc/syntrop/routerd.toml
Restart=on-failure
RestartSec=3s
WatchdogSec=30s
Slice=ai.slice
# Memory constraints
MemoryHigh=24M
MemoryMax=32M
# Security & Sandboxing
User=syntrop
Group=syntrop
NoNewPrivileges=yes
ProtectSystem=strict
ProtectHome=yes
PrivateTmp=yes
ProtectKernelModules=yes
ProtectKernelTunables=yes
RestrictAddressFamilies=AF_UNIX AF_INET AF_INET6
RestrictNamespaces=yes
MemoryDenyWriteExecute=yes
RestrictRealtime=yes
RestrictSUIDSGID=yes
LockPersonality=yes
# cgroup v2 & systemd-oomd Protection
ManagedOOMPreference=avoid
OOMScoreAdjust=-800
OOMPolicy=stop