git
ask
hub
Privacy
Terms
Sign in with GitHub
rankun203/llm-proxy
↗
Stars ·
1
Language ·
Python
Simplify and Visualize This Repo
Scan the safety of this repo
Self Host this repo
Check out more work of Developer: rankun203
Ask anything about this repo to start.
Set up llm-proxy-ondemand in front of my vLLM server so the model starts on first request and shuts down after 30 minutes of inactivity. Walk me through the configuration steps.
I have access to a university SLURM cluster and want to run a large language model on-demand using llm-proxy-ondemand. Help me configure passwordless SSH and the reverse SSH tunnel so requests from my laptop reach the model on the cluster node.
Write a script that sends a chat completion request to my llm-proxy-ondemand endpoint using the OpenAI Python client, so I can test that the proxy starts the model server on first request and forwards the response.
Help me change the idle timeout in llm-proxy-ondemand from the default 30 minutes to 10 minutes, and explain where that setting lives in the config.
Full explanation on explaingit →
📎
Send
By chatting or signing in you agree to the
Terms
and chat-message logging (revocable in
History
).