What is a BYOK AI router?

Bring your own key means your own provider account supplies model access. A router connects your coding client to that account and translates supported API formats. RouteMyAgent hosts that router for you: you rent a routing slot, attach your NVIDIA API key, and connect your existing agent.

You are renting routing software access. NVIDIA supplies the model inference. A routing slot is not a dedicated GPU, a virtual machine, or a bundle of inference credits.

Where your request goes

  1. Your locally installed coding client sends a request to your RouteMyAgent endpoint using a rental-bound router token.
  2. The router checks the selected model, active rental, attached provider key and applicable request limits.
  3. The router sends the supported request to NVIDIA using your key and returns the response in the client’s API format.
  4. Your coding client executes its local tools and sends their results back for the next model turn.

RouteMyAgent supports OpenAI-compatible Chat Completions and Responses routes, plus an Anthropic-compatible Messages adapter. This makes API translation possible; it does not turn NVIDIA models into OpenAI or Anthropic models or grant access to those companies’ hosted products.

Two credentials, different jobs

Your NVIDIA API key authorizes inference with NVIDIA. Attach it privately in the signed-in portal. The application encrypts the attachment in a separate credential store for restart recovery. Normal ledger backups exclude this store. Disconnecting it or rental expiry removes the saved attachment. Temporary memory copies can remain; this is not a zero-knowledge service.

Your router token authorizes your coding client to use the current rental. Enter that token privately when your separate rental launcher asks. Keep both credentials out of chat prompts, repositories, screenshots and support messages. Read the privacy notice for how prompts, operational records and provider copies are handled.

What the trial and hourly rental include

Comparing the total cost of AI coding? Our affordable AI coding guide separates provider, client and routing charges, with hourly examples and alternatives to renting a router.

The first trial provides 30 minutes of routing, once per account, subject to capacity. Setup and idle time are paused. Prepare your authorized NVIDIA key and installed client before sending model requests. Signing in alone does not consume routing time.

Paid rentals use prepaid routing credit. Choose the model tier and review the current rate in the pricing section before confirming. The current tiers are GLM 5.3 and Kimi K3; each rental permits its selected model only. Provider pricing, quotas and availability remain separate. There is no automatic renewal.

New rentals buy active routing use in whole-hour blocks upfront. Setup and idle time are paused; overlapping requests count once. Failed provider requests do not consume usage time, but cancellation after returned output or exhaustion of the budget still consumes use already delivered. Access stops when usage time runs out, followed by a five-minute renewal hold for the same model before the slot is released. A restart preserves purchased usage time and restores the encrypted key for an active rental. Older wall-clock rentals retain their original deadline and five-minute renewal hold. Purchased credit and rental time are non-refundable except where required by law; review the service terms before buying.

Prepare your client, then verify it

Use the guide for Claude Code or Claude Desktop, or Codex and OpenCode. During an active rental the portal generates setup instructions for the exact endpoint, model and installed client version. These instructions preserve your normal configuration and create a separate rental launcher with an Undo script.

A browser chat can help write setup files but cannot configure your computer without local execution access. A local coding agent can apply the instructions, then test a small text request and a harmless tool-result round trip. Those tests use your NVIDIA quota. A successful website sign-in is not a client connection test.

Rate limits and interrupted responses

A rented slot does not bypass NVIDIA limits. On a 429 response, honor Retry-After and reduce parallel work. If a request times out or ends without a complete response, it may already have reached the provider. Inspect the error and local task state before resubmitting. More retries do not repair an invalid model, expired rental or missing key.

Prepare your first connection →