vLLM

vLLM is a Large Language Model (LLM) runtime designed for high performance and scalability. It is the industry standard from single-GPU to large-cluster deployments.

Important

The vLLM LMOD Module is the ONLY approved method of running LLMs on ARC Resources

Supported Resources

LLMs require specialized hardware to run effectively; only certain resources are supported

Cluster (Partition)

Accelerator

Supported

Orion

None

No

Hercules (v1)

None

No

Hercules (v2)

AMX (CPU)

Yes

Ptolemy

CUDA (GPU)

Yes

Atlas

CUDA (GPU)

Yes

Morrill

CUDA (GPU)

Yes

Warning

Do not run on Orion or Hercules v1 - they lack AMX instructions and will have extremely poor performance.

Additionally, ARC Workstations equipped with NVIDIA GPUs with 8GB of VRAM or greater are also supported, however, they will run a much smaller model. It is not recommended to use agents with this smaller model, performance may be subpar.

Setup

vLLM is provided as a module on desktops and supported clusters. Load it into your environment with

module load vllm

vLLM also requires a few user-owned files to run; examples of these files have been provided. To populate a user-owned directory with these example files, run the following

setup-vllm-example.sh

By default, it will populate a directory named offline-vllm within the current one. Optionally, you can specify where the files are placed by specifying it as an argument on the commandline.


Operation modes

vLLM has two operating modes, “Online”, and “Batch” Mode.

Online Mode

In Online mode, vLLM acts as an LLM provider. The application provides an OpenAI compatible REST API. Users can point any OpenAI compatible client software at the endpoint without issue.

Note

“Online” Mode does not mean that vLLM is operating via the web, it’s simply their naming scheme for the REST API mode.

To launch vLLM in Online mode, run

sbatch --account <your_account> vllm-api.sbatch # on clusters
# or
bash vllm-api.sbatch                            # on desktops

Note

If you just ran the example setup script. you’ll need to provide your account either by modifying the sbatch script or by passing the --account <your_account> argument to the sbatch commandline

Please note that startup may take quite a while as vLLM prepares the necessary kernels and loads the models.

Once running, point clients at the provided host with the provided API key the submission script outputs.

Batch Mode

In Batch Mode, vLLM acts as a python module. Users can inference with vLLM however they see fit in code. This can be very useful if the user has a large amount of inputs they need to inference on. The examples provide a simple python example that takes a JSON file of inputs as the prompt for inference, and outputs a JSON of responses.

First, modify inputs.json with your desired prompts. Each conversation is an array of messages with role and content:

[
  [
    {"role": "user", "content": "What is the capital of France?"}
  ],
  [
    {"role": "user", "content": "Explain quantum computing."}
  ]
]

Then start the job with

sbatch --account <your_account> vllm-offline.sbatch # on clusters
# or
bash vllm-offline.sbatch                            # on desktops

The outputs will be recorded into outputs.json