model format = MLX (specific for macOS)
The .gguf file extension stands for "GPT Graph Unified Format", which is a model file format introduced by the ggml project for storing and running large language models (LLMs) like LLaMA, Mistral, and others in a unified, efficient way.
🔧 Key Features of .gguf: Self-contained format: Includes model weights, metadata, vocabulary/tokenizer, and tensor layout information in a single file.
Cross-model support: Standardizes support for models like LLaMA, Falcon, Mistral, GPT-J, GPT-NeoX, and others.
Optimized for inference: Works with ggml-based runtimes such as:
llama.cpp
gpt4all
text-generation-webui
Supports quantization: GGUF files can be quantized (e.g., Q4_K_M, Q8_0) to reduce size and improve performance on CPUs.