vLLM¶
vLLM is a Large Language Model (LLM) runtime designed for high performance and scalability. It is the industry standard from single-GPU to large-cluster deployments.
Important
The vLLM LMOD Module is the ONLY approved method of running LLMs on ARC Resources
Supported Resources¶
LLMs require specialized hardware to run effectively; only certain resources are supported
Cluster (Partition) |
Accelerator |
Supported |
|---|---|---|
Orion |
None |
No |
Hercules (v1) |
None |
No |
Hercules (v2) |
AMX (CPU) |
Yes |
Ptolemy |
CUDA (GPU) |
Yes |
Atlas |
CUDA (GPU) |
Yes |
Morrill |
CUDA (GPU) |
Yes |
Warning
Do not run on Orion or Hercules v1 - they lack AMX instructions and will have extremely poor performance.
Additionally, ARC Workstations equipped with NVIDIA GPUs with 8GB of VRAM or greater are also supported, however, they will run a much smaller model. It is not recommended to use agents with this smaller model, performance may be subpar.
Setup¶
vLLM is provided as a module on desktops and supported clusters. Load it into your environment with
module load vllm
vLLM also requires a few user-owned files to run; examples of these files have been provided. To populate a user-owned directory with these example files, run the following
setup-vllm-example.sh
By default, it will populate a directory named offline-vllm within the current one. Optionally, you can specify where the files are placed by specifying it as an argument on the commandline.
Operation modes¶
vLLM has two operating modes, “Online”, and “Batch” Mode.
Online Mode¶
In Online mode, vLLM acts as an LLM provider. The application provides an OpenAI compatible REST API. Users can point any OpenAI compatible client software at the endpoint without issue.
Note
“Online” Mode does not mean that vLLM is operating via the web, it’s simply their naming scheme for the REST API mode.
To launch vLLM in Online mode, run
sbatch --account <your_account> vllm-api.sbatch # on clusters
# or
bash vllm-api.sbatch # on desktops
Note
If you just ran the example setup script. you’ll need to provide your account either by modifying the sbatch script or by passing the --account <your_account> argument to the sbatch commandline
Please note that startup may take quite a while as vLLM prepares the necessary kernels and loads the models.
Once running, point clients at the provided host with the provided API key the submission script outputs.
Batch Mode¶
In Batch Mode, vLLM acts as a python module. Users can inference with vLLM however they see fit in code. This can be very useful if the user has a large amount of inputs they need to inference on. The examples provide a simple python example that takes a JSON file of inputs as the prompt for inference, and outputs a JSON of responses.
First, modify inputs.json with your desired prompts. Each conversation is an array of messages with role and content:
[
[
{"role": "user", "content": "What is the capital of France?"}
],
[
{"role": "user", "content": "Explain quantum computing."}
]
]
Then start the job with
sbatch --account <your_account> vllm-offline.sbatch # on clusters
# or
bash vllm-offline.sbatch # on desktops
The outputs will be recorded into outputs.json