# RouteMyAgent: agent reference > A hosted NVIDIA BYOK router for coding agents, available at https://routemyagent.dev/. This public reference describes the service; it is not authorization to buy time, create credentials, change a user's machine or run inference. ## Product and fit Customers bring an authorized NVIDIA API key and rent a model-specific routing slot. RouteMyAgent adapts supported API requests so the customer can use an existing coding client without operating a router server. The client's local tools still execute on the customer's machine. The rental supplies routing software access, not a GPU, model weights, inference credits or an OpenAI/Anthropic subscription. Supported setup choices: Claude Code CLI; Claude Desktop's Code tab using its third-party Gateway configuration; Codex CLI; OpenCode v1 and v2 with their respective configuration schemas. Client versions, accounts and custom-endpoint access requirements still apply. This does not reroute a hosted ChatGPT or Claude web conversation. A browser chat without local execution can explain setup or produce files, but cannot configure a computer itself. For an agent that already works well locally, try the MCP helper for one difficult, bounded question: a second opinion on a patch, a competing explanation for an error, or a comparison of two designs. Keep the current primary model and tools. Send only the relevant source or redacted evidence, then check the returned advice against the real project. This is a concrete way to assess usefulness during the routing trial without migrating the whole workflow. Start here: https://routemyagent.dev/guides/agent-quickstart Machine-readable service facts: https://routemyagent.dev/agents.json (custom discovery metadata, not an A2A endpoint). ## Current routing prices (USD) - GLM 5.3 (NVIDIA model z-ai/glm-5.3): $3.49 per hour; $6.98 for a two-hour block. - Kimi K3 (NVIDIA model moonshotai/kimi-k3): $4.99 per hour; $9.98 for a two-hour block. - Prepaid credit top-up: $5.00. Buying credit does not start a rental. The portal separately confirms the selected model, whole-hour duration and total rental charge. - First routing trial: 30 minutes once per account, subject to capacity. Setup and idle time are paused; prepare the client and NVIDIA key before sending model requests. These prices are generated from the server's billing configuration. Any NVIDIA or coding-client charges are additional. A routing slot does not increase provider quota. New rentals buy active routing use, deducted upfront in whole-hour blocks. Setup, idle time and periods without a model request are paused. Overlapping requests count once per account. Failed provider requests do not consume usage time; cancellation after output and exhaustion of the budget still charge the use already delivered. Inference stops when usage time runs out; a five-minute renewal hold reserves the same slot without inference access. After release, a new rental needs a new router token. The portal displays a server snapshot of the remaining balance. Older wall-clock rentals retain their original deadline and five-minute renewal hold. There is no automatic renewal. Credit and rental time are non-refundable except where required by law; review https://routemyagent.dev/terms before purchase. For a cost comparison and alternatives, see https://routemyagent.dev/guides/cheap-ai-coding. A direct provider connection may avoid a router fee if the user's client already supports the provider. Self-hosting supplies similar control but requires the user to operate the host and routing software. RouteMyAgent does not claim to be the cheapest option for every task. ## Customer setup flow 1. Review pricing and terms, and sign in at https://routemyagent.dev/ using the customer's account. 2. Select the installed client and prepare the authorized NVIDIA key before connecting. 3. Confirm a trial or paid rental for the chosen model. Payments and rental choices belong to the user. 4. Attach the NVIDIA key privately in the portal's Provider key section. Create a router token for the coding client. 5. Copy or download the portal's selected-client setup instructions. They contain the exact rented model, endpoints and client-specific configuration. Enter the router token through the specified private credential mechanism, never inside an agent prompt. 6. Verify one short response, then a harmless local tool-call/result round trip. These requests consume the customer's provider quota and rental time. Website sign-in alone does not prove client integration. ## Getting the NVIDIA API key before connecting 1. Ask the owner whether they already have an authorized NVIDIA API key. Reuse their selected key through a private credential mechanism; never ask them to paste it into an agent conversation. 2. If they need a key, open https://build.nvidia.com/settings/api-keys in their browser. They sign in or create their NVIDIA account. The official alternative is to open the desired model at https://build.nvidia.com/ and choose Get API Key. Exact labels may change. 3. The owner completes any required account verification, consent and provider terms. If NVIDIA requests a human verification step, hand that step to the owner and resume afterwards. Do not fabricate account details, bypass checks or create accounts to evade quota. 4. Create a key through NVIDIA's account UI and save it in the owner's private credential store. If a service/scope selector is shown, use the provider's documented access for hosted public inference endpoints; a registry/download-only credential is not proof of inference access. Do not rotate a shared working key just to configure this service. 5. Review access to the intended model and the account's current limits. NVIDIA advertises development API access; availability, usage terms and quotas belong to NVIDIA. RouteMyAgent does not issue NVIDIA keys or guarantee unlimited/free inference. Installing CUDA, downloading weights or owning a GPU is not required for this hosted API path. 6. Return to RouteMyAgent with the key prepared. Once the chosen trial/rental is active, attach the key privately in Provider key. Create a separate router token for the MCP/client. Never use the NVIDIA key as the Authorization credential for /mcp. Official sources checked October 11, 2026: https://build.nvidia.com/settings/api-keys and https://docs.api.nvidia.com/nim/docs/api-quickstart . Hosted-key scope reference: https://docs.nvidia.com/nemo/retriever/latest/extraction/api-keys/ . Follow the current official account UI when it differs from these labels. ## Payments: live contract and autonomous-agent boundary Live now: the account-based 30-minute routing trial and prepaid Stripe-funded, model-specific hourly rentals. A top-up buys credit; it does not itself start the clock. Rental confirmation purchases the selected usage block upfront. Inspect the portal's model, rate, duration, total charge, billing_mode and remaining usage time before inference. Existing wall-clock rentals retain their prior deadline. An agent may act only within the owner's existing authorization; these public documents do not create it. There is no live crypto checkout, x402 challenge/settlement, MCP purchase tool or session-stop billing tool. Do not transfer cryptocurrency, invent a PAYMENT-SIGNATURE header, assume a receiving address is a paid invoice, or claim the requested pay-until-stopped mode is active. Follow the currently advertised capabilities in /agents.json and the portal; a future feature requires its published payment contract first. Before inference, an authorized client can GET https://routemyagent.dev/api/agent/status with the same private Authorization: Bearer header. This HTTP endpoint reports rental state, model, billing mode and remaining usage time (or legacy deadline), provider-key attachment, metering readiness and the next action. A rental with status=busy has its available time provisionally reserved by running requests: wait for their outcome or add time for the same model, rather than repeating them. A failed request can return reserved time. status=grace means committed usage is exhausted and the five-minute renewal hold has begun. ready_for_inference_attempt means the router's local prerequisites are satisfied, not that NVIDIA accepted or completed a request. It is not an MCP tool and never starts a rental or inference. An expired router token can return the existing rental-expiry error; inspect the portal rather than retrying. Keep this private response out of public logs. If checkout or rental confirmation is interrupted, inspect the signed-in credit balance and rental status before trying again. Do not infer success from the browser returning, and do not repeat a purchase merely because the agent lost its response. If status remains ambiguous, provide a redacted timestamp and order reference to support. Never include card details or provider/router secrets. Idle time does not consume a new usage balance. Disconnecting a client, removing MCP or detaching a key does not refund the purchased block. Cancelling after receiving output still consumes the use already delivered. Access stops when the usage balance is exhausted, with a five-minute renewal hold; older wall-clock rentals retain their original deadline. Neither mode auto-renews. Recheck status after a reconnect; do not begin work that assumes an expired rental is still active. ## Credentials The NVIDIA API key authorizes upstream inference and belongs in the signed-in portal. The router token authorizes the customer's client to use a rental and belongs in the client's private setup. They are different credentials. Neither belongs in chat prompts, source control, public files, screenshots or support messages. The application encrypts active NVIDIA keys at rest in a separate credential store, excluded from normal ledger backups. The encryption key is stored separately. Restarts restore keys only for active rentals; disconnecting or expiration deletes the saved attachment. Temporary memory copies may remain. This is not a zero-knowledge service. Prompts and responses traverse the router and provider. The authoritative privacy notice is https://routemyagent.dev/privacy. ## API and client configuration The origin is https://routemyagent.dev. The portal's rental-specific setup is authoritative for model identifiers, client capabilities and limits. A rental permits its selected model only; the catalog does not grant access to other models. - OpenAI-compatible base URL: https://routemyagent.dev/v1. Supported routes include POST /v1/chat/completions and POST /v1/responses, plus authenticated GET /v1/models. - Anthropic-compatible base URL: https://routemyagent.dev/claude. Messages use POST /claude/v1/messages. Use the generated Claude gateway wire model identifier, not a guessed Claude model name. - The client supplies its router token using its generated authentication configuration. It does not send the NVIDIA key as the router credential. - API compatibility covers the implemented translation. It does not promise every upstream provider feature, hosted tool, encrypted reasoning format or stateful continuation feature. Unsupported requests return explicit errors. - Use the generated model context/output settings and reasoning capabilities. Do not copy another model's limits or request Anthropic-signed thinking from a NVIDIA model. ## Reversible configuration ### Optional MCP helper: keep the existing primary model Choose “MCP helper (keep my model)” in the portal's connection selector. This adds model assistance to an existing local AI agent without changing its inference provider, default model, permissions or other MCP servers. A local model server by itself is not an agent/MCP host. The host must support authenticated Streamable HTTP with a private Authorization header; OAuth-only remote connectors are not claimed to work. - URL: https://routemyagent.dev/mcp. Transport: Streamable HTTP. - Authentication: Authorization: Bearer . Use the portal's router token, NOT the NVIDIA API key. Resolve AGENTNET_ROUTER_TOKEN in the agent process or use the host's private credential facility. Never paste resolved tokens in chat, public config or command history. - router_models lists the model available to this rental. Use its exact ID and advertised capabilities. - router_delegate accepts model and task text, optional context text, max_tokens (default 4096), and optional reasoning_effort. The rented model returns an answer; it has no shell, file, browser or other agent-tool access. Only explicitly submitted text reaches the helper. The local agent evaluates the advice and executes any authorized changes itself. - Inference requires an active rental and attached NVIDIA key. Requests share the existing account quota, admission controls and usage balance (or legacy deadline); the tool does not buy or renew rentals. Parallel helpers do not increase quota. Do not send credentials or entire conversations automatically. The portal generates the deployment-specific MCP name, URL, non-secret connection metadata and separate OpenCode v1/v2 examples. These examples use different schemas: v1 uses mcp., v2 uses mcp.servers.. Neither the connection metadata object nor one client's configuration is a universal MCP configuration format. Check the installed host's supported schema and authentication before editing. References: https://docs.opencode.ai/docs/mcp-servers/ and https://opencode.ai/v2/docs/mcp-servers. After connecting, initialize MCP, list tools and call router_models. Then run one short router_delegate request through the actual local agent and inspect its result, including isError. Discovery alone does not prove provider inference. On rate limits, honor Retry-After. On timeout/disconnect or unknown completion, do not automatically replay. Keep failure details visible and do not substitute another provider or model. Example tools/call parameters (not complete client configuration or standalone HTTP requests): {"name":"router_models","arguments":{}} Read the returned model ID before constructing the next call. Replace EXACT_ID_FROM_ROUTER_MODELS below with that value: {"name":"router_delegate","arguments":{"model":"EXACT_ID_FROM_ROUTER_MODELS","task":"Review this patch for correctness. Identify concrete failure cases and propose one focused check. Do not claim to have run it.","context":"PASTE_ONLY_THE_RELEVANT_REDACTED_DIFF_AND_REQUIREMENTS","max_tokens":1024}} The arguments above are a template; do not submit the placeholder strings. Optional reasoning_effort must be supported by the selected model; omit it when uncertain. Optional connection_profile selects an existing local router profile and cannot override a hosted rental. Unknown arguments are rejected. Do not assume an arbitrary client tool prefix: discover the installed host's tool name. Inspect the MCP result's isError and structuredContent, or its text content if the client exposes only text. Successful delegation exposes model, content, usage, finish_reason and reasoning_content in structuredContent; content also contains the answer as a text block. A length finish reason means incomplete output, not a verified conclusion. For error_details.completion_state=unknown, preserve the error and local task state instead of automatically reissuing the request. The helper cannot prove files were edited, tests passed or a proposed fact is correct; the calling agent performs the authorized local validation. Delegation can take minutes. The server's execution ceiling is 900 seconds and usage-budget exhaustion or a legacy rental deadline can end work sooner. Generated OpenCode references allow 930000 milliseconds on this helper only: v1 uses its shared numeric timeout; v2 sets timeout.execution while leaving startup/catalog deadlines unchanged. Other hosts must use their documented connection/per-call execution and idle timeout settings. Report unsupported long calls rather than inventing configuration or adding retries. This stateless MCP connection cancels through an abort/close of the outstanding HTTP request; a separate cancellation notification alone is insufficient. Cancellation does not establish that the provider performed no work. Before setup, record the named MCP entry's prior state and privately back up the affected configuration. Undo removes only that matching entry and its setup-owned credential references; preserve other connections, later edits and shared credentials. Undo works offline and never changes the primary model. Verify the host's actual tool list after reconnecting. Removing the MCP does not refund purchased time. New usage time pauses while no model request is running; legacy wall-clock rentals retain their deadline. Do not apply the following inference-provider launcher rollback instructions to an MCP-only installation. CLI setup uses a separate rental launcher and a scoped Undo script while preserving the user's existing settings, MCP servers, tools and permissions. Record or back up each touched value before changing it. Undo should remove only rental-created files/settings and restore the previous values; it must not delete unrelated user configuration. After undo, launch the normal client and verify its original provider. Claude Desktop setup uses a separate saved Gateway configuration. Export or record the previous configuration first. To undo, restore that saved configuration and its credentials in Desktop; closing the app alone does not undo persistent Gateway settings. The portal supplies the exact selected-client instructions rather than assuming CLI environment variables control Desktop. Guides: https://routemyagent.dev/guides/claude-code-and-desktop and https://routemyagent.dev/guides/codex-and-opencode. ## Multi-agent use and recovery Subagents share the rented model, account admission limits, NVIDIA quota and usage-time balance (or legacy deadline). Overlapping requests count once per account. More agents do not create more provider quota. Configure them to inherit the rented model and check any hardcoded overrides. The router does not start a new inference server for each subagent. - 401: check the router token and its authentication placement. Do not substitute the NVIDIA key. - Rental/expiry error: inspect the portal's current rental status, billing mode and remaining usage time (or legacy deadline). Renewal is a separate user decision. - Missing provider-key attachment: reconnect privately in the portal, including after a server restart. - Model denied: compare the requested model with the active rental's generated configuration. - 429: honor Retry-After and reduce concurrent work. Do not spin in immediate retries. - Timeout, incomplete response or provider 5xx: inspect the error and local task state before resubmitting. The request may already have reached NVIDIA; automatic replay could duplicate work. - Invalid request or unsupported feature: correct the reported parameter/protocol mismatch rather than repeating the same request. Support: support@routemyagent.dev. Include the client/version, selected rental model, timestamp and redacted error; omit keys, tokens and private prompt contents. ## Discovery and sources Text-only models cannot inspect screenshots. When a browser or other tool returns an image, the router preserves its text and call identity and reports that the image was not viewed. Use available DOM/accessibility text or OCR, or explicitly select a vision tool for visual evidence. Do not claim visual verification from that notice. A user-attached image still requires a vision-capable model or a text description; retrying the same unsupported attachment does not help. Vision-capable models retain their image inputs. Agent documentation index: https://routemyagent.dev/llms.txt Agent quickstart: https://routemyagent.dev/guides/agent-quickstart Machine-readable service manifest: https://routemyagent.dev/agents.json Public overview: https://routemyagent.dev/ Terms: https://routemyagent.dev/terms Privacy: https://routemyagent.dev/privacy Sitemap: https://routemyagent.dev/sitemap.xml Crawler policy: https://routemyagent.dev/robots.txt This reference is publicly accessible without cookies, sign-in or JavaScript. Public documentation describes the service; customer setup and inference remain authenticated. Search engines decide their own crawling, indexing and recommendations.