Files
wild-directory/vllm/versions/0
Paul Payne 4d983819c9 feat(supabase): add services and statefulset for database management
feat(synapse): update ingress to use traefik ingress class and bump version

feat(syncthing-discovery): introduce syncthing discovery service with deployment and ingress

feat(syncthing-relay): add syncthing relay server with deployment and ingress configuration

fix(taiga): update liveness and readiness probes to use tcpSocket for health checks

fix(taiga): change PVC access mode to ReadWriteMany for media and static storage

feat(traefik): add icon and ignore rules for traefik service

docs(ushahidi): add notes for Redis configuration and Laravel startup probe adjustments

feat(ushahidi): implement dedicated Redis deployment for Ushahidi

fix(vllm): update deployment strategy and readiness/liveness probes for improved stability

fix(writefreely): pin writefreely image version to v0.15.1 for consistency

docs(zulip): add notes for TLS-terminating reverse proxy configuration and expected behavior
2026-07-02 21:34:27 +00:00
..

vLLM

vLLM is a fast and easy-to-use library for LLM inference and serving with an OpenAI-compatible API. Use it to run large language models on your own hardware.

Dependencies

None, but requires a GPU node in your cluster.

Configuration

Key settings configured through your instance's config.yaml:

  • model - Hugging Face model to serve (default: Qwen/Qwen2.5-7B-Instruct)
  • maxModelLen - Maximum sequence length (default: 8192)
  • gpuProduct - Required GPU type (default: RTX 4090)
  • gpuCount - Number of GPUs to use (default: 1)
  • gpuMemoryUtilization - Fraction of GPU memory to use (default: 0.9)
  • domain - Where the API will be accessible (default: vllm.{your-cloud-domain})

Access

After deployment, the OpenAI-compatible API will be available at:

  • https://vllm.{your-cloud-domain}/v1

Other apps on the cluster (such as Open WebUI) can connect internally at http://vllm-service.llm.svc.cluster.local:8000/v1.

Hardware Requirements

This app requires a GPU node in your cluster. Adjust the gpuProduct, gpuCount, and memory settings to match your available hardware.