Ankk98/bhandaar

On-disk cache with streaming and cross-process concurrency

★ 0Forks 0RustGitHub ↗Compare

Project website ↗

cachecratefile-storageobject-storagerust

README

bhandaar (भंडार)

On-disk LRU/TTL file cache for Rust with cross-process concurrency via flock(2).

bhandaar (भंडार) — Hindi for "storehouse, repository"

License: MIT Rust

Features

  • LRU eviction — least recently used entries are evicted first
  • TTL support — per-entry time-to-live with lazy expiration
  • Cross-process safe — uses filock for flock(2) locking
  • Sharded design — user-defined shard count for concurrent access
  • Greedy storage — fills up to max size, evicts only when needed
  • Streaming API — handle large files without loading into memory
  • Two-level hashing — SHA-256-based directory nesting, supports millions of entries

Quick Start

use bhandaar::{Bhandaar, BhandaarConfig};
use std::time::Duration;

fn main() -> Result<(), Box<dyn std::error::Error>> {
    let cache = Bhandaar::new(BhandaarConfig {
        root: "/tmp/bhandaar".into(),
        max_size: 1024 * 1024 * 1024, // 1 GB
        shards: 16,
        default_ttl: None,
    })?;

    // Put data
    cache.put("user:123:avatar", &image_data)?;

    // Put with TTL (expires in 1 hour)
    cache.put_with_ttl("session:abc", &session_data, Duration::from_secs(3600))?;

    // Get data (loads into memory)
    if let Some(data) = cache.get("user:123:avatar")? {
        println!("Got {} bytes", data.len());
    }

    // Stream large files (nothing loaded into memory)
    if let Some(file) = cache.get_file("large:file")? {
        // file is a File handle — read it however you want
        let mut reader = std::io::BufReader::new(file);
        // ... process line by line, etc.
    }

    Ok(())
}

API

In-Memory API (for small data)

// Put
cache.put("key", b"data")?;
cache.put_with_ttl("key", b"data", Duration::from_secs(3600))?;

// Get (loads entire file into memory)
let data = cache.get("key")?;

// Check existence
let exists = cache.contains("key")?;

// Remove
let removed = cache.remove("key")?;

// Get all keys
let keys = cache.keys()?;

// Clear all entries
cache.clear()?;

// Get current size
let size = cache.size();

Streaming API (for large files)

// Get a file handle — nothing loaded into memory
if let Some(file) = cache.get_file("large:file")? {
    // Read line by line
    use std::io::{BufRead, BufReader};
    let reader = BufReader::new(file);
    for line in reader.lines() {
        let line = line?;
        // Process line...
    }
}

// Stream data into a writer
use std::io::Write;
let mut writer = std::io::BufWriter::new(std::fs::File::create("output.txt")?);
cache.get_file("large:file")?
    .map(|mut f| std::io::copy(&mut f, &mut writer));

// Put from a reader (streaming write)
let reader = std::io::Cursor::new(b"large data...");
cache.put_reader("key", reader, 1024)?; // size must be known

// Put from a reader (auto-detect size)
let reader = std::io::Cursor::new(b"data...");
let actual_size = cache.put_reader_auto_size("key", reader, None)?;

Error Handling

use bhandaar::BhandaarError;

match cache.put("key", b"data") {
    Ok(()) => println!("Stored"),
    Err(BhandaarError::AlreadyExists(key)) => println!("Key exists: {}", key),
    Err(BhandaarError::CapacityExceeded) => println!("Cache full"),
    Err(e) => println!("Error: {}", e),
}

S3 Cache Example

Use bhandaar as a local disk cache for S3 objects:

// See examples/s3_cache.rs for full implementation

struct S3Cache {
    cache: Bhandaar,
    s3_client: Box<dyn S3Client>,
}

impl S3Cache {
    fn get_object(&self, bucket: &str, key: &str) -> Result<Option<File>, Error> {
        let cache_key = format!("{}/{}", bucket, key);

        // Check cache first (returns File handle, not Vec<u8>)
        if let Some(file) = self.cache.get_file(&cache_key)? {
            return Ok(Some(file));
        }

        // Cache miss — fetch from S3 and stream to disk
        self.fetch_and_cache(bucket, key, &cache_key)?;
        Ok(self.cache.get_file(&cache_key)?)
    }
}

Architecture

Directory Structure

Files are stored using two-level SHA-256 hashing:

cache_root/
├── a1/
│   ├── b2/
│   │   ├── c3d4e5f6g7h8.dat       # Cached file
│   │   ├── c3d4e5f6g7h8.dat.meta  # Sidecar metadata (bincode)
│   │   └── ...
│   └── ...
└── ...
  • Max ~65K directories at each level
  • <10K files per leaf directory
  • Supports millions of entries

Concurrency Model

  • Per-shard locking — each shard has its own Mutex for in-process coordination
  • filock integration — per-file flock(2) for cross-process safety
  • Atomic size tracking — lock-free total size counter

Eviction Strategy

  • LRU — least recently used entries evicted first
  • TTL — expired entries lazily deleted on access
  • Greedy — fills up to max size, evicts only when needed

Streaming Design

  • get_file() — returns std::fs::File handle for streaming reads
  • put_reader() — streams data from Read trait to disk
  • put_reader_auto_size() — streams without knowing size upfront
  • Zero memory overhead — large files never loaded entirely into RAM

License

MIT — see LICENSE for details.

Contributors

Ankk98

Issues