npx skills add ...
npx skills add huggingface/prime-rl --skill inference-server
npx skills add huggingface/prime-rl --skill inference-server
Start and test the prime-rl inference server. Use when asked to run inference, start vLLM, test a model, or launch the inference server.
Always use the inference entry point — never vllm serve or python -m vllm.entrypoints.openai.api_server directly. The entry point runs setup_vllm_env() which configures environment variables (LoRA, multiprocessing) before vLLM is imported.
The inference entrypoint supports optional SLURM scheduling, following the same patterns as SFT and RL.
Each node runs an independent vLLM instance. No cross-node parallelism — TP and DP must fit within a single node's GPUs.
Add dry_run = true to generate the sbatch script without submitting:
The server extends vLLM with:
/v1/chat/completions/tokens — accepts token IDs as prompt input (used by multi-turn RL rollouts)/update_weights — hot-reload model weights from the trainer/load_lora_adapter — load LoRA adapters at runtime/init_broadcaster — initialize weight broadcast for distributed trainingsrc/prime_rl/entrypoints/inference.py — entrypoint with local/SLURM routingsrc/prime_rl/inference/server.py — vLLM env setupsrc/prime_rl/configs/inference.py — InferenceConfig and all sub-configssrc/prime_rl/inference/vllm/server.py — FastAPI routes and vLLM monkey-patchessrc/prime_rl/templates/inference.sbatch.j2 — SLURM template (handles both single and multi-node)configs/debug/infer.toml — minimal debug config# inference_slurm.toml
output_dir = "/shared/outputs/my-inference"
[model]
name = "Qwen/Qwen3-8B"
[parallel]
tp = 8
[slurm]
job_name = "my-inference"
partition = "cluster"uv run inference @ inference_slurm.toml# inference_multinode.toml
output_dir = "/shared/outputs/my-inference"
[model]
name = "PrimeIntellect/INTELLECT-3-RL-600"
[parallel]
tp = 8
dp = 1
[deployment]
type = "multi_node"
num_nodes = 4
gpus_per_node = 8
[slurm]
job_name = "my-inference"
partition = "cluster"uv run inference @ config.toml --dry-run truecurl http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "Qwen/Qwen3-0.6B",
"messages": [{"role": "user", "content": "Hi"}],
"max_tokens": 50
}'