exllama logo

Exllama

Exllama is a memory-efficient implementation of Hugging Face transformers for LLaMA models using quantized weights, enabling high-performance NLP on modern GPUs with minimal memory use.

exllama homepage

Key Features

  • Quantized Weights

    Uses 4-bit GPTQ quantization to minimize memory footprint.

  • Processor Affinity

    Configurable CPU affinity to optimize hardware resource usage.

  • Flexible Stop Conditions

    Allows custom stopping criteria during content generation.

  • Web UI Support

    Includes a simple web interface for interactive model use.

Get Started

(0)

Share & Save

Share on Social Media

Why Choose Exllama

  • Memory Efficiency:

    Reduces GPU memory usage with 4-bit quantized weights for large LLaMA models.
  • High Performance:

    Optimizes inference speed on modern NVIDIA GPUs including RTX series.
  • Sharded Models:

    Supports loading and running sharded model files for scalability.

Pricing

Exllama is an open-source project available for free on GitHub. No pricing plans apply.

About Exllama

Exllama is a memory-efficient implementation of Hugging Face transformers for LLaMA models using quantized weights, enabling high-performance NLP on modern GPUs with minimal memory use.

What Exllama Does

Exllama enables efficient deployment of LLaMA transformer models by using quantized 4-bit weights to reduce memory usage while maintaining high inference performance on modern GPUs.

It supports sharded model loading, configurable processor affinity for hardware optimization, and flexible stop conditions for content generation, allowing users to tailor performance to their hardware setup.

Ideal for NLP developers, machine learning engineers, and researchers, Exllama facilitates experimentation and deployment of large language models in resource-constrained environments and supports integration with Python projects and web UIs.

Try Exllama

Pros & Cons

  • Open Source

    Freely available with community contributions on GitHub.

  • Modern GPU Support

    Optimized for NVIDIA RTX 30-series and newer GPUs.

  • Limited GPU Compatibility

    Older GPUs with poor FP16 support may perform poorly.

  • Work in Progress

    Some features and optimizations are still under development.

Frequently Asked Questions

What GPUs does Exllama support?

Exllama is optimized for NVIDIA RTX 30-series and newer GPUs with good FP16 support.

Is Exllama free to use?

Yes, Exllama is an open-source project available for free on GitHub.

How do I install Exllama?

Clone the GitHub repo, install dependencies via pip, and run the provided Python scripts.

Does Exllama support sharded models?

Yes, it supports loading multiple sharded .safetensors files for large models.

Is there a user interface for Exllama?

A simple web UI is included, which can be run locally using Python and Flask.

Similar Tools You Might Like

Discover more AI-powered tools that complement your workflow

Visit Tool Page

List Your AI Tool & Reach Thousands of Users

Join 500+ AI innovators already thriving on our platform. Get visibility, feedback, and boost your conversions.

Expand Your Audience

Connect with over 50,000 AI enthusiasts actively looking for tools like yours.

Boost Your Authority

Get verified reviews and ratings to build credibility in the AI marketplace.

Drive Conversions

Our premium placements and targeted audience deliver quality leads and sign-ups.