GitHub - skegdb/skeg: A multi-tenant vector database focused on extreme RAM efficiency. Lightweight, scalable, and optimized for high-density deployments.

GitHub

6 min read Original article ↗

skeg

The vector database that fits.
Multi-tenant, disk-first, RAM-frugal. Recall 1.0 at a fraction of the memory.

crates.io release CI MSRV 1.88 Apache-2.0 benchmarks

Skeg - The memory-efficient vector DB with high recall. | Product Hunt


skeg stores the full vectors on SSD and keeps only a small quantized working set in RAM. That trade buys recall 1.0 on a memory footprint the RAM-resident engines cannot reach, which is what matters when memory is the contested resource: thousands of tenants on one box, or a vector store sharing a machine with the model it serves.

Key-value and vectors live in the same engine, behind a Redis-compatible wire protocol.

Quickstart

docker run -d --name skeg -p 6379:6379 -v skeg-data:/var/lib/skeg \
  --entrypoint /usr/local/bin/skeg-resp3 ghcr.io/skegdb/skeg:latest

Any Redis client talks to it:

$ redis-cli -3 -p 6379
> SET greeting "hello"
OK
> SKEG.VINDEX.CREATE docs 1024 tq2 disk
OK
> SKEG.VSET docs 1 <1024-float vector as bytes>
OK
> SKEG.VSEARCH docs 10 100 <query vector bytes>
1) "1"
2) (double) 0.987

Vector operations sit under SKEG.* so they stay clear of the Redis command surface. The native protocol runs on 7379 and is the default entrypoint; skeg-resp3 above serves RESP3 on 6379. Full walkthrough, command reference and filter grammar: docs/getting-started.md.

Which protocol

Use RESP3 for application integrations. It is the supported public API and names the vector tiers directly: f32, int8, tq1, tq2, tq4, binary.

The native transport on 7379 exists for specialised clients. It is versioned, and the version decides which tiers it can name:

v1 v2
kinds 0=f32 1=int8 2=binary the same, plus 3=tq1 4=tq2 5=tq4
kind 3 rejected: historical clients used it for PQ tq1

A v2 client opens with NativeHello (op 0x84) and reads the tier capability mask it gets back. v1 byte meanings are unchanged, so an existing client keeps working.

Benchmarks

Reproducible from skeg-bench (public harness, real embeddings, brute-force ground truth). Measured single-machine on Apple Silicon; the RAM ratios are hardware-independent.

Single-tenant, 100K x 1024-dim, recall against exact brute force. Every engine at a reasonable default, with LanceDB tuned to recall 1.0 for a fair fight:

engine serve RAM recall@10 p50 latency
skeg (tq2) 47 MB 1.000 2.5 ms
Milvus Lite 108 MB 0.934 2.7 ms
LanceDB (IVF-PQ) 198 MB 0.998 59 ms
hnswlib (raw HNSW) 426 MB 0.985 2.0 ms
Chroma (HNSW) 682 MB 0.985 3.9 ms
Qdrant (HNSW, f32) 885 MB 0.997 2.6 ms

Co-resident with a model: a 3B LLM answering RAG over 1M vectors, both on one M1 Pro (16 GiB). The index stays on SSD and the resident set stays flat.

Co-resident, 1M vectors backend RSS p50 backend RSS max
skeg (pq128) 54 MiB 67 MiB
Qdrant (HNSW) 254 MiB 2,387 MiB

Backend RSS while a 3B LLM serves RAG, swept from 10K to 1M vectors on an M1 Pro 16 GiB. skeg stays under 80 MiB; Qdrant climbs into multi-GiB territory.

The full matrix, plus the multi-tenant and container-OOM runs, is on the dashboard.

What it does not win

skeg is not the lowest-latency single-query engine: Qdrant is comparable on p99 and raw hnswlib is faster. One process saturates near 780 QPS at 1024-dim, past which you scale out with processes. Cold bulk-loads rebuild the index.

Multi-tenancy

Tenancy is a property of the storage layout rather than a filter convention. Each tenant gets its own index, so a query has no physical path to another tenant's vectors, and there is no filter to misconfigure. An adversarial leak-fuzz holds it to that: query one tenant's index with another tenant's exact vector and zero rows cross the boundary, every time.

On top of that isolation:

  • Hard quotas: max_vectors and max_disk_bytes, set and read at runtime through SKEG.QUOTA.SET / SKEG.QUOTA.GET.
  • Fair eviction, so a noisy tenant cannot starve a quiet one out of the cache.
  • Authentication via HELLO 3 AUTH user pass (argon2id), with prefix-routed namespaces.

Details in docs/multi-tenancy.md.

Install

Docker

docker run -d --name skeg -p 7379:7379 -v skeg-data:/var/lib/skeg \
  ghcr.io/skegdb/skeg:latest

The image carries both binaries and publishes for linux/amd64 and linux/arm64. The default entrypoint is skeg on 7379; for RESP3 override with --entrypoint /usr/local/bin/skeg-resp3 and publish 6379. An Ollama companion setup lives in docker-compose.example.yml.

Homebrew (macOS and Linux ARM)

brew tap skegdb/tap
brew install skeg

Installs both binaries and a launchd/systemd service.

Pre-built tarball

TARGET=aarch64-apple-darwin   # see Platforms below for the full list
TAG=$(curl -s https://api.github.com/repos/skegdb/skeg/releases/latest | grep tag_name | cut -d'"' -f4)
curl -L -o skeg.tar.gz \
  "https://github.com/skegdb/skeg/releases/latest/download/skeg-${TAG}-${TARGET}.tar.gz"
tar -xzf skeg.tar.gz && ./skeg --help

Each tarball ships a .sha256 next to it. Pin a version from the releases page.

From crates.io

cargo install skeg-server

Builds from source into $CARGO_HOME/bin. Needs a Rust toolchain (MSRV 1.88).

From git

git clone https://github.com/skegdb/skeg
cd skeg
cargo build --release --bin skeg --bin skeg-resp3

Binaries land in target/release/.

Platforms

target tarball container
aarch64-apple-darwin yes none
aarch64-unknown-linux-gnu yes linux/arm64
x86_64-unknown-linux-gnu yes linux/amd64
x86_64-unknown-linux-gnu, AVX-512 -avx512 suffix :<version>-avx512

Kernel selection happens at runtime, after a CPU feature check, so a binary is never tied to the machine that built it. On aarch64 the NEON kernels are always compiled in. On x86_64 the AVX2 kernels are, and one binary covers every x86_64 CPU with a scalar fallback below AVX2.

The AVX-512 kernels are the exception: they are compiled in only when asked for, because they earn their place only where AVX-512 offers an instruction AVX2 lacks: VNNI, VPOPCNTDQ, mask registers, a 16-entry vpermps table. Where a kernel is the same technique at twice the width it loses on cores that split 512-bit operations, so it ships built and tested but not selected. Building them needs Rust 1.89, one release above the MSRV, which is why they are a separate artifact rather than the default.

An AVX-512 build still runs on a CPU without AVX-512; the extra kernels simply never get selected. To build one yourself:

cargo build --release --bin skeg --bin skeg-resp3 --features skeg-server/avx512

cargo test -p skeg-simd --test coverage prints which kernel runs on which instruction set, and why any gap is a gap.

Documentation

Guides in docs/:

Long-form design and benchmark write-ups are on the project blog: Constraints as Method, Seven More Hypotheses, The Substrate, What Was Measured.

Published crates: skeg-proto, skeg-simd, skeg-platform, skeg-telemetry, skeg-resp3, skeg-core, skeg-vector, skeg-server, skeg-tenant, skeg-server-tenant, skeg-multi-tenant. Network adapters live in skeg-rigging and skeg-rigging-net.

Contributing

Bug reports, design discussions, and pull requests are welcome. Before opening a PR run cargo fmt, cargo clippy --workspace --all-targets -- -D warnings, and cargo test --workspace. The pre-push hook at .githooks/pre-push runs the same three; enable it with git config core.hooksPath .githooks (a docs-only push can skip it with SKIP_PREPUSH=1).

Security

Report security issues by opening an issue with a brief description and a request to take the conversation private. See SECURITY.md.

License

Apache-2.0. See NOTICE for attribution.