Skip to content

nihilistau/shannon-prime-system-engine

Folders and files

NameName
Last commit message
Last commit date

Latest commit

Β 

History

921 Commits
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 

Repository files navigation

type reference
title shannon-prime-system-engine β€” the inference engine + resident daemon + memory agency
description Repo README for the Shannon-Prime ENGINE β€” the sp-daemon server, CUDA backend, served Gemma-4-12B chat over the resident persistent-KV cache, memory agency, the L5-cosine + attribute-gate recall/faithfulness stack, NIGHTSHIFT curator, and KAIROS heartbeat. States repo scope vs the other four repos, the served-chat data path, exact env flags + gate names, and build/run one-liners.
tags
shannon-prime
engine
sp-daemon
gemma4
recall
faithfulness
adr-002
cuda
l5-cosine
attr-gate
timestamp 2026-07-01 00:00:00 UTC
resource shannon-prime-system-engine/README.md
sp_status GREEN-LIVE
sp_gate none
sp_commit fc2e846
sp_repro none

shannon-prime-system-engine β€” the inference engine + resident daemon + memory agency

This is THE ENGINE of Shannon-Prime. It is the sp-daemon server, the CUDA backend (wire_cuda_backend), the served Gemma-4-12B chat over a resident persistent-KV cache, the model-owned memory agency (forget / decide / merge), the recall + faithfulness stack (L5-cosine recall + a deterministic attribute-grounding gate), NIGHTSHIFT curator, KAIROS heartbeat, Telepathy delegate wiring, and EAGLE/MTP draft. It consumes the math core (shannon-prime-system, carried here as the lib/shannon-prime-system submodule) through the frozen L1 C ABI, and is orchestrated by the harness (shannon-prime-harness).

License: MIT (LICENSE). GitHub: nihilistau/shannon-prime-system-engine.

Read this first for the whole system: lattice papers/PPT-LAT-KEYSTONE.md. This README is this repo's scope, data path, flags, and gates. Public site: Position Is Arithmetic.

What Shannon-Prime is

A fully local, byte-exact, auditable language-model organism: it serves Google's Gemma-4-12B on a single RTX 2060, through our own inference engine, on an exact-integer arithmetic substrate (O_K, dual-prime negacyclic CRT-NTT). It owns a working memory β€” it learns facts from conversation, recalls them faithfully, declines when it does not know, and forgets / supersedes / merges on its own judgement. No cloud, no third-party inference, no telemetry. Every mechanism is a default-off env flag = a strict no-op / byte-identical null floor when unset; every number has a reproducing command and a gate.

What this repo controls vs the other four repos

Repo Role This repo talks to it via
shannon-prime-system-engine (here) THE ENGINE. sp-daemon server + /v1/chat loop, accelerator backends, served 12B chat over the resident persistent-KV cache, memory agency, recall/faithfulness stack, NIGHTSHIFT, KAIROS. β€”
shannon-prime-system The math core β€” decode loop, O_K / CRT-NTT, exact-islands, ARM recall router. Frozen. the L1 C ABI (sp_session_register_*_backend), carried as lib/shannon-prime-system submodule
shannon-prime-harness Orchestration β€” the agency scheduler (run_agency.py), corpus/curator tooling, eval drivers. drives the daemon over HTTP + shells the gates
shannon-prime-lattice Docs / canon / STATE β€” papers, contracts, ADRs, scoreboard, ledgers. Source of truth for status. pointers only (no code dependency)
shannon-prime (public face) Position_Is_Arithmetic β€” receipts-first public papers + LEDGER.md. citations only

Architecture map

   shannon-prime-harness            shannon-prime-lattice (docs / canon / STATE)
   run_agency.py, gates,            PPT-LAT-KEYSTONE, VERIFIED-SCOREBOARD,
   curator/eval drivers             ADR-002, FINDINGS-LEDGER, START-HERE
        β”‚ HTTP / shell                        β–² pointers only
        β–Ό                                     β”‚
  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
  β”‚  ENGINE (this repo)  β€”  tools/sp_daemon/  (Rust, HTTP / SSE / WebSocket)   β”‚
  β”‚  /v1/chat loop (routes.rs) Β· recall+faithfulness (recall.rs) Β·             β”‚
  β”‚  memory agency FORGET/DECIDE/MERGE Β· NIGHTSHIFT curator Β· KAIROS heartbeat β”‚
  β”‚  Β· Telepathy delegate Β· EAGLE/MTP draft Β· sampler/tokenizer/sessions      β”‚
  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                   β”‚  frozen L1 C ABI
                                   β”‚  (sp_session_register_forward_backend,
                                   β”‚   sp_session_register_kvdecode_backend, …)
        β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
        β–Ό              β–Ό                       β–Ό               β–Ό
   CUDA backend    CPU backend            Vulkan backend   Hexagon HVX
  (wire_cuda_       (overlay dispatch      (desktop)        (edge)
   backend:          into math-core
   gemma4 fwd/       decode)
   decode, OK_Q4B
   dp4a, SP_BYTEEXACT)
        β”‚
        └──────────► lib/shannon-prime-system  (MATH CORE, submodule:
                     decode loop, O_K / CRT-NTT, exact-islands, L1 ABI)

The served-chat data path (/v1/chat)

The daemon holds a resident persistent-KV cache and drives the 12B token-by-token over the L1 ABI. A turn flows:

HTTP POST /v1/chat  (tools/sp_daemon/src/server.rs β†’ routes.rs)
   β”‚
   β”œβ”€ apply chat template (tokenizer.rs)
   β”œβ”€ PERSISTENT O(1) KV  (SP_PERSIST_KV, default-ON): if the committed cache is a
   β”‚     strict prefix of this turn, extend it instead of re-prefilling β€” flat VRAM,
   β”‚     O(1) context growth  (routes.rs: kv:: prefill/decode/rewind)
   β”‚
   β”œβ”€ DECIDE stage (latent, ADR-002):  choose whether to recall / decline / forget /
   β”‚     supersede β€” deciders read latent, they do NOT execute:
   β”‚        Β· SP_RECALL_L5      β€” L5-cosine paraphrase recall selector (see below)
   β”‚        Β· SP_RECALL_ATTR_GATE β€” attribute-grounding gate β†’ zero-inference decline
   β”‚        Β· SP_FORGET / SP_DECIDE β€” memory agency writers
   β”‚
   └─ EXECUTE stage (clean text): synthesize the answer in a fresh clean context
         from the DECIDE result β€” NEVER fusing the latent decision into generation.

Governing law — ADR-002 Decide→Execute spine: decide in latent, execute in clean text, never fuse; deciders don't execute. (lattice papers/PPT-LAT-ADR-002-DECIDE-EXECUTE-SPINE.md.)

Recall + faithfulness stack (CLOSED end-to-end, 2026-07-01)

The faithfulness axis is closed on the served chat. Honest tiers: PROVEN / GREEN-LIVE (gated and on the served path), PARKED (built, gated, deliberately not on the hot path), HONEST-NEGATIVE.

Stage What it does Tier Flag(s) Gate / commit
L5-cosine recall selector matches the live query's global-layer-5 embedding against each episode's stored L5 query-key by cosine; on a confident match (Ο„=0.30) delivers the episode TEXT in-context GREEN-LIVE SP_RECALL_L5=1 (SP_RECALL_L5_TAU, default 0.30) G-L5-RECALL-LIVE (86.89% paraphrase LIVE) Β· d9099cd
Attribute-grounding gate β†’ zero-inference decline closes the zero-prior / private-data hole (SNE crucible: 80% confab, 5% secret-leak on novel entities). If the query's salient words are ABSENT from the fact (attr_absent_ratio β‰₯ Ο„) and the query carries a high-entropy entity token (query_has_entity_token), it emits a deterministic symbolic decline and skips the gemma4 forward entirely β€” confabulation/leak becomes mathematically impossible; paraphrase recall untouched GREEN-LIVE SP_RECALL_ATTR_GATE=1 (SP_RECALL_ATTR_TAU, default 0.5) G-SNE-ATTRGATE-ZEROINF (confabβ†’0, leakβ†’0, recall 100%) Β· HEAD fc2e846
Generative recall/reject judge 12B reads candidate memory TEXTS and picks the one that answers the query, or NULL PARKED SP_B3_JUDGE hard-foreign kill-test: 0 benefit vs L5-direct+Ο„ (PASSed 15/18) β€” parked, not on the hot path

Relevant code: tools/sp_daemon/src/recall.rs (l5_query_embed, cos512, attr_absent_ratio, query_has_entity_token), tools/sp_daemon/src/routes.rs (the SP_RECALL_L5 branch + the symbolic_decline synthesis-seam short-circuit).

Other engine mechanisms (all default-off = byte-identical null floor)

Mechanism Tier Flag Gate / commit
Persistent O(1) KV β€” resident cache, flat VRAM, O(1) context growth PROVEN / default-ON SP_PERSIST_KV (unset β‡’ on; =0 forces flat) G-PERSIST-KV
Byte-exact exact-integer forward β€” RMSNorm/softmax/GELU/RoPE + attention as exact-integer device kernels = exact arithmetic / cross-machine determinism / AUDITABILITY, not compression gated-GREEN / default-off SP_BYTEEXACT G-BYTEEXACT-FORWARD-12B (OFF 4.6665 byte-identical / ON 4.6569 parity / run-to-run bit-identical) Β· 69c0588
Model-owned memory agency β€” STORE (NIGHTSHIFT) + FORGET + DECIDE/MERGE (supersede a changed fact, or consolidate two complementary facts) GREEN-LIVE SP_FORGET, SP_DECIDE G-FORGET / G-DECIDE / G-MERGE
NIGHTSHIFT offline curator β€” 12B model-call secret-extractor β†’ teacher-forced causal-ablation admit (TAU=βˆ’8) β†’ conformant MEM-OKF emit gated-GREEN / live PENDING SP_NIGHTSHIFT_OFFLINE=1 G-NIGHTSHIFT-CURATOR criteria 1-4 Β· 6107f3e
Autonomous W_c recall head β€” learned latent selector, (E+1)-way NULL argmax + bounded-mass replay gated-GREEN (superseded on the hot path by L5-cosine) SP_B3_WC G-CHAT-B3-WC-DEPLOY / -DIV2 Β· edc8079

The boundary thesis runs through all of it: O_K wins on exact arithmetic (the container); every structure-on-content lever (MΓΆbius, split-prime Dirichlet carriers, entropy-coded Frobenius codes, T2-MΓΆbius on the real embedding) is measured-inert and kept as an honest negative.

SP-SWARM β€” the private memory mesh (tools/sp_swarm, L0–L4 GREEN, default-off)

The sp_swarm Rust crate lives in this repo. It is replication + discovery over the content-addressed MEM-OKF store β€” not a new store β€” so the operator's own nodes converge on one signed memory. Five gated layers, each reusing a proven engine asset:

Layer Role Reuses Gate
L1 content addressing addr = sha256(norm(body))[:16] + the C2-SimHash episode address class the byte-exact sha2 addressing + recall::Projection C2 sigs G-SWARM-RUST-PARITY (Rust↔Python byte-parity, incl. the CRLF-normalize fix)
L2 replication have/want diff β†’ verified pull β†’ converge; verify-on-arrival before write the MEM-OKF store layout G-SWARM-REPLICATE-CONVERGE
L3 provenance Ed25519 sign-on-write / verify-vs-roster (invite-only) ed25519-dalek (audited; ↔ pynacl parity) G-SWARM-PROVENANCE-ED25519
L0 transport QUIC (TLS 1.3) + Ed25519 mutual roster handshake; LIST/GET/SIM/DONE wire network::quic_shard (quinn 0.11 / rustls 0.23-ring) G-SWARM-TRANSPORT-QUIC
L4 discovery C2-SimHash SIM shortlist gossip β†’ exact-fetch verify recall::agree (256-bit Hamming) G-SWARM-C2-INDEX, G-SWARM-GOSSIP-DISCOVERY

Integrated into sp-daemon behind an optional, default-off swarm feature (SP_SWARM=1; G-SWARM-NODE / G-SWARM-DAEMON-WIRE), plus a standalone headless sp-swarm-node binary. Honest negative (G-SWARM-C2-SEMANTIC): C2-256 is a shortlist (recall@5 0.885 == L5-cosine top-1), not a top-1 oracle (0.607) β€” used as a hint, confirmed by exact-fetch; bit-count is the lever. It ships curated records, never raw latent (ADR-002). Remaining = multi-host deployment only.

:: build the daemon WITH the mesh (default-off until SP_SWARM=1)
cargo build --release --features wire_cuda_backend,swarm --target-dir target-wirecuda --bin sp-daemon
:: the swarm crate's own gates (deps cached β‡’ works offline)
cd tools\sp_swarm && cargo test --features transport

Env flags: SP_SWARM (=1 spawns), SP_SWARM_PORT (7777), SP_SWARM_KEY (Ed25519 seed file), SP_SWARM_ROSTER (node_id <pubkey_hex> lines), SP_SWARM_NODE_ID, SP_SWARM_PEERS, SP_SWARM_INTERVAL_S, SP_SWARM_ROOT/SP_OKF_ROOT. Full call surface + wire protocol + the 9 gates: lattice papers/PPT-LAT-MESH-API.md; design + rejected mechanics: papers/PPT-LAT-DESIGN-SWARM-MEMORY-MESH.md.

Build

Two distinct toolchains (do not mix). The CPU / math-core build uses clang-cl (MSVC-ABI β€” lib/shannon-prime-system/core/exact_islands/exact_islands.c uses __int128, which cl.exe cannot compile). Authoritative from-clean chain + every gotcha: lattice papers/BUILD-ENV-TOOLCHAIN.md (gate G-CLEAN-BUILD).

Backend Toolchain Build dir
CUDA (canonical) VS2019 BuildTools + CUDA, ninja, sm_75 (dev RTX 2060) build-cuda/
CPU (canonical) clang-cl (MSVC-ABI) β€” NOT cl.exe build-cpu/
:: build the served daemon (drives the real 12B end-to-end over the L1 ABI)
cargo build --release --features wire_cuda_backend --target-dir target-wirecuda --bin sp-daemon

:: a CUDA gate
cd build-cuda && ninja test_gemma4_cuda && tests\test_gemma4_cuda.exe

:: the daemon decode gate (32/32 tokens == oracle, VRAM flat O(1))
cargo run --release --features wire_cuda_backend --bin sp_wire_cuda_decode_gate

GPU numbers need warmup + a long window + both clocks pinned. Git on these repos: native PowerShell, not the Linux mount (the mount CRLF-churns large files + locks).

Run the live organism

run_console.bat               :: coherent + byte-exact + O(1) 12B chat -> http://127.0.0.1:3000/
run_console_recall.bat        :: + the autonomous librarian
run_console_nightshift.bat    :: + the offline NIGHTSHIFT curator (SP_NIGHTSHIFT_OFFLINE=1)

The faithfulness stack is turned on with SP_RECALL_L5=1 SP_RECALL_ATTR_GATE=1. Alongside the daemon, run the harness agency scheduler for the heartbeat consolidation + maintenance (shannon-prime-harness/run_agency.py). State a fact and it is captured; ask an attribute the fact does not state and it declines in microseconds with no forward instead of confabulating; say "forget X" and it is dropped; state a contradicting/complementary fact and DECIDE supersedes or MERGEs it. Full procedure: lattice papers/PPT-LAT-KEYSTONE.md Β§10.

Key engine surfaces

  • tools/sp_daemon/ β€” the resident daemon (feature wire_cuda_backend drives the real 12B):
    • src/routes.rs β€” the /v1/chat loop: persistent O(1) KV (SP_PERSIST_KV), the SP_RECALL_L5 branch + symbolic_decline synthesis-seam short-circuit, the memory agency (SP_FORGET, SP_DECIDE/MERGE), the EOT bias, the SP_CURRENT_CONVO consolidation hook.
    • src/recall.rs β€” the recall/faithfulness primitives (l5_query_embed, cos512, attr_absent_ratio, query_has_entity_token; plus the learned WcHead).
    • src/nightshift_curator.rs β€” the offline curator (opens its own kvdecode handle, never touches the served cache). src/kairos.rs / kairos_runner.rs β€” the heartbeat control-plane.
    • src/telepathy.rs, src/eagle_accept.rs β€” delegate wiring + EAGLE/MTP draft-accept.
  • src/backends/cuda/cuda_forward.cu β€” the gemma4 CUDA forward + decode (per-layer SWA/global geometry, shared-KV, AltUp/PL=0, softcap); the OK_Q4B k_gemv_q4b_dp4a kernel; the CUDA-graph decode path; the SP_BYTEEXACT exact-integer islands; the additive gemma4_kv_decode_logits (the daemon's token-by-token decode entry) + the persistent-KV ABI.
  • src/backends/cpu/ β€” overlay dispatch into the math-core decode. vulkan/, hexagon/ are desktop / edge targets.
  • tools/sp_transcode/ β€” sp_transcode --st: the safetensors-direct pipeline, the ONLY trusted gemma4-12B weight path (the GGUF lane is dead). Writes OK_Q8 / OK_Q4B .sp-model.
  • tools/sp_dsp_smoke/ β€” the L2 universal Rust crate: dual-prime Barrett / mod-q matmul / Garner CRT / NTT ladder, bit-exact-gated; the 4 islands' exact-integer references in src/sp_islands_q_ref.rs.
  • tools/sp_swarm/ β€” the SP-SWARM private memory mesh crate (L0–L4): lib.rs (addressing + have/want replication), similarity.rs (C2Index discovery), transport.rs (QUIC + Ed25519 roster + serve/pull_from/discover_similar/run_node), bin/sp_swarm_node.rs. Optional default-off dep of the daemon (--features swarm). Gates in tests/ + tests/fixtures/swarm/. See the SP-SWARM section above + lattice papers/PPT-LAT-MESH-API.md.
  • Gates / receipts β€” tests/test_gemma4_cuda.c, tests/test_xbar_p1_cuda.c, tests/fixtures/.

Canon / status pointers (lattice repo β€” source of truth)

  • papers/PPT-LAT-KEYSTONE.md β€” the whole system, current + complete.
  • papers/VERIFIED-SCOREBOARD.md β€” box-by-box honest tiers.
  • papers/PPT-LAT-ADR-002-DECIDE-EXECUTE-SPINE.md β€” the governing decideβ†’execute law.
  • papers/PPT-LAT-FINDINGS-LEDGER.md β€” the findings ledger.
  • START-HERE.md β€” orientation. papers/PPT-LAT-STATE.md β€” the PROVEN record.

Non-negotiables (receipts-first)

  • No number without a reproducing command + a gate/commit. Bit-exact-when-off: every SP_* overlay is a strict no-op by default β€” verify the null floor before claiming a delta.
  • Honest tiers, exact words. GREEN-LIVE = gated AND on the served path by default. PARKED / gated-GREEN / default-off = passes a gate but is NOT on the hot path. HONEST-NEGATIVE stays attached (the diffusion-judge falsification, the boundary-thesis inert levers, the parked judge).
  • No silent gate revision β€” surface upstream.
  • Check the code + commits + git fetch before trusting memory. lib/shannon-prime-system can diverge from the standalone math-core checkout β€” check behind before building/committing.
  • Served-model misbehavior is almost always ours (template / decode / sampler / forward / prompt), not the weights β€” verify vs llama.cpp + our PPL first.

About

Clean from-scratch inference engine for shannon-prime-lattice. NTT-based attention, two-node CRT-sharded inference path, KSTE-encoded KV state.

Topics

Resources

License

Stars

1 star

Watchers

0 watching

Forks

Packages

 
 
 

Contributors