Keep your agent. Add a second opinion.

Your local agent already knows the project. Give it a focused outside review through MCP: a tricky diff, a debugging hypothesis, or a choice between two designs. It keeps its model, tools and files; the rented NVIDIA model sees only the text you choose to send.

A useful first task: ask for concrete failure cases in a small patch and one check you can run locally. Compare the advice with the source and run that check. Evaluate a real result before expanding the workflow.

Prepare your 30-minute routing trial → Service manifest for agents

The trial is once per account, subject to capacity. Prepare the key and client first: setup time counts. Provider quotas and client access requirements still apply. Current pricing · Terms · Privacy

1. Get your NVIDIA API key first

  1. If you already have an authorized NVIDIA API key, use your private credential store. Do not paste it into the agent's chat.
  2. Open NVIDIA API-key settings and sign in or create your NVIDIA account. Alternatively, open your intended model in NVIDIA's catalog and choose Get API Key.
  3. Complete NVIDIA's required account verification and review its terms. An agent can explain the steps, but the account owner handles any verification or consent the provider requires. Do not bypass verification or create extra accounts to evade quotas.
  4. Create the key in NVIDIA's UI and store it privately. If scope choices appear, follow NVIDIA's instructions for hosted public inference access. A registry-only key is not proof of inference access. Preserve shared working keys rather than rotating them unnecessarily.
  5. Check access to the model you want and the account's current limits. No GPU, CUDA installation or model download is needed for this hosted API path.

NVIDIA controls eligibility, terms, quotas and availability. Its advertised development API access is separate from the routing trial; RouteMyAgent does not issue NVIDIA keys or promise unlimited inference. Official instructions: API Catalog quickstart and key authentication reference. Checked October 11, 2026; follow the current official UI if labels change.

2. Check your agent host before starting the clock

You need an agent host that supports Streamable HTTP MCP with a private Authorization header. A bare model API is not an MCP host. OAuth-only connectors are not claimed to work. Choose MCP helper (keep my model) in the portal to get instructions for your installed client; do not paste another client's JSON schema into yours.

Record the named MCP entry's existing state and privately back up the affected settings. Keep your primary model, unrelated MCP servers and tool permissions unchanged. Use the host's supported secret facility for the router token, never a resolved secret in a prompt, repository or command history.

3. Confirm the rental and attach the right credential

  1. Sign in with your own account, review the selected model and prepare the client and NVIDIA key.
  2. Confirm the eligible trial or the paid rental's rate, duration and total. Buying prepaid credit alone does not start a rental.
  3. In the signed-in portal's Provider key section, attach the NVIDIA API key privately. Create a separate router token for your agent host.
  4. Use the portal's generated MCP connection settings: URL https://routemyagent.dev/mcp, transport Streamable HTTP, private Authorization: Bearer <router token>.

The NVIDIA key goes to the portal, not the MCP Authorization header. The attachment is encrypted at rest for restart recovery and decrypted in memory for requests. Disconnect or expiry deletes it; a restart restores only active-rental keys. This is not zero-knowledge hosting: the router and provider handle the text you send. Read the privacy notice.

Preflight without running a model: an authorized client can GET https://routemyagent.dev/api/agent/status with that private router-token header. Read the rental status, model, billing mode and remaining usage time (or legacy deadline), key attachment and next action. ready_for_inference_attempt confirms local prerequisites only, not provider availability or a successful response. A busy rental means running requests have temporarily reserved the available time: wait for their outcome or add time, without duplicating pending work. A failed request can return the reservation. A grace status means committed usage is exhausted and the renewal hold has begun. This is an HTTP endpoint, not an MCP tool, and does not start a rental. If the token has expired, inspect the portal's rental status.

4. Discover, delegate once, verify

Initialize MCP and list the server's tools through your host. Call router_models first. These examples are tools/call parameters, not client configuration or complete HTTP requests:

{"name":"router_models","arguments":{}}

Use the returned exact model ID in the next call. Replace both placeholders with your selected model and relevant redacted context:

{
  "name": "router_delegate",
  "arguments": {
    "model": "EXACT_ID_FROM_ROUTER_MODELS",
    "task": "Review this patch for correctness. Identify concrete failure cases and propose one focused check. Do not claim to have run it.",
    "context": "RELEVANT_REDACTED_DIFF_AND_REQUIREMENTS",
    "max_tokens": 1024
  }
}

Only model and task are required. context is optional; max_tokens defaults to 4096. Supply reasoning_effort only when supported by the model. The optional connection_profile is for existing local router profiles; it cannot change hosted rental access. Unknown fields are rejected.

Inspect isError and structuredContent (or the exposed text content). Successful delegation exposes model, content, usage, finish_reason and reasoning_content in structured content; the answer is also returned as a text block. A length finish reason means output was truncated. Your agent must verify suggestions itself: the helper has no filesystem, shell or browser access and cannot execute tests.

5. Understand payment and failure states

Live payment flow: the account-based trial and prepaid Stripe-funded hourly rentals. MCP does not buy or renew rentals. There is currently no crypto checkout, x402 settlement or session-stop billing tool. Do not send cryptocurrency expecting access, invent payment headers, or treat discovery metadata as permission to spend.

New purchases buy whole-hour blocks of active routing use, paid upfront. Setup and idle time are paused. Overlapping model requests count once per account. Failed provider requests do not consume usage time; cancellation after output has been returned and exhaustion of the time budget still consume the use already delivered. Access stops when usage time runs out; a five-minute hold lets you add time for the same model before the slot is released. Adding more requires a separate choice. Removing MCP does not refund purchased time. Older wall-clock rentals retain their original deadline and five-minute renewal hold.

Calls may take minutes. The helper's execution ceiling is 900 seconds; usage exhaustion or a legacy rental deadline can stop them earlier. Configure your host's documented execution timeout appropriately. Cancel by aborting the outstanding HTTP request; a standalone cancellation notification is insufficient for this stateless connection. Cancellation does not prove that NVIDIA performed no work.

6. Undo cleanly

Remove only the matching MCP entry and setup-owned credential references, preserving later edits and other connections. Reconnect the host and check its tool list. Your original primary model remains selected. Revoke an unused router token through the portal if you no longer need it; do not delete a shared NVIDIA key. Undo works offline and does not refund purchased time. New usage time pauses when no request is running; older wall-clock rentals keep their existing deadline.

For full provider routing instead of an MCP helper, use the separate guides for Claude Code / Desktop or Codex / OpenCode, which describe their own configuration rollback.

Try one useful task →