Skip to content

Testnet Troubleshooting Guide ​

Tier B (cargo budget-report without --replay) deploys a contract, funds an account through friendbot as part of that deploy, and simulates every exported function. Every one of those three steps can fail for reasons that have nothing to do with your contract's code. This page catalogs the failures you're likely to hit, what they mean, and what to do about each one — so a testnet blip doesn't send you searching your own contract for a bug that isn't there.

INFO

Every entry below is labelled either reproduced — the exact wording quoted was produced by actually triggering the failure — or from source — the wording comes directly from the string literal in cargo-budget-report/src/main.rs or live.rs, quoted rather than paraphrased, but not independently triggered against a live testnet from this environment. Both are accurate; the label just tells you which kind of confidence to expect.

First: is it hung, or is it retrying? ​

Deploy, invoke-build, and the simulateTransaction RPC call are each wrapped in the same retry loop (run_with_retry in main.rs). By default that's up to 4 attempts, with the delay before each retry doubling: 2s → 4s → 8s. Unless --quiet is passed, every retry prints a line to stderr naming the attempt and the reason it is retrying, then a spinner ticks for the length of the backoff so the wait does not look like a hang:

Deploy attempt 1/4 failed: Error: friendbot rate-limited (try again later). Retrying in 2 s...
⠹ backing off 2 s before the next attempt
Deploy attempt 2/4 failed: Error: friendbot rate-limited (try again later). Retrying in 4 s...
⠹ backing off 4 s before the next attempt
Deploy attempt 3/4 failed: Error: friendbot rate-limited (try again later). Retrying in 8 s...

(from source — the format string in run_with_retry and the backoff_sleep spinner. The spinner is drawn only on an interactive stderr; under --quiet, a redirect, or --retry-backoff-secs 0 it is a plain sleep.)

With the default settings, the worst case for a single call site is 2 + 4 + 8 = 14 seconds of sleeping before it gives up — bounded and predictable, not a hang. If your terminal has sat silent for longer than that with no new output and no exit (and no spinner), something other than the documented retry loop is going on (network still connecting, DNS still resolving) — that's a genuine hang, not this mechanism.

To change this behavior:

  • --max-retry-attempts N (or [retry].max_attempts in budget.toml) — 1 disables retry entirely, useful when you want a fast, clear failure instead of a slow one.
  • --retry-backoff-secs N (or [retry].initial_backoff_secs) — changes the initial delay.

See the full retry policy reference for the precedence between these and defaults.

Not every failure is retried ​

The retry loop only re-attempts failures it classifies as plausibly transient. Everything else fails on the first try, on purpose — retrying a deterministic failure (a typo'd contract ID, a malformed argument) four times with growing delays would only make the eventual, unavoidable failure slower. The exact classifier (is_transient_error in main.rs) retries a message if it contains any of these substrings, case-insensitively:

rate limit, rate-limited, ratelimit, 429, too many requests, connection, timed out, timeout, reset by peer, broken pipe, 503, 502, unavailable, temporarily, try again

Anything else — including a missing contract, a bad XDR, or an RPC-reported simulation error — is treated as permanent and fails immediately. This is the practical way to tell the two failure classes apart: if you saw retry lines in the output, the tool already decided this looked transient; if it failed on attempt 1 with no retry lines, it decided the failure was deterministic — which almost always means something in your configuration, not the network, needs fixing.

When a deploy does give up, a second classifier (deploy_diagnostics::classify in main.rs) reads the final error and picks the guidance to print — rate limiting, a service outage, an unreachable network, or an unfunded source account. Only the first three are worth retrying; an unfunded account is called out explicitly as something that will not resolve by waiting.

Your configuration vs. a bad network day ​

The single fastest way to tell which kind of problem you have:

You see...It usually means
Several attempt N/4 failed. Retrying in ... s lines, then final failureThe network (or friendbot) is having a bad day. Retrying later, or re-running, often just works.
Failure on the very first attempt, no retry lines at allSomething in your setup is wrong in a way retrying cannot fix: a missing binary, an unfunded/misconfigured identity, a bad path, a bad argument.

Use the table below to go from the specific symptom to the specific cause.

Failure catalog ​

Friendbot rate limiting ​

Symptom: Deploy fails, retries a few times with growing delays, then either succeeds or exhausts all 4 attempts with a final message like:

Failed to deploy amm-pool-contract after 4 attempts.
Cause: rate limiting. Friendbot / the RPC is throttling requests from this IP.
  What to do: wait ~60 seconds and re-run — the limit is per-IP and lifts on its
  own. Running from CI on a shared runner makes this more likely; a short sleep
  before the step usually clears it.
Last error: stellar contract deploy failed: Error: friendbot rate-limited (try again later)

(from source — deploy_contract_with_retry combined with deploy_diagnostics::classify picking the RateLimited guidance, and the friendbot wording the stellar CLI itself returns, which this project's own test fixture at cargo-budget-report/tests/fixtures/fake_bin/stellar reproduces for MOCK_STELLAR_FAIL_COUNT tests.)

Cause: Testnet's friendbot service rate-limits funding requests. Deploying a fresh (unfunded) source identity triggers a friendbot call as part of stellar contract deploy; under load, or when many CI jobs hit it in a short window, that call is throttled.

What to do: This is the classic "bad network day" case — it's why the retry loop exists. Let the retries run; if all 4 still fail, wait a minute and re-run the command. If this happens constantly in CI, fund your source identity once ahead of time (stellar keys generate alice --network testnet --fund, or top it up manually) so deploy doesn't need friendbot on every run, and consider raising --max-retry-attempts for that job.

Friendbot / account not yet confirmed on-ledger ​

Symptom: Same shape as rate limiting — deploy fails and retries — but the underlying cause is different: friendbot's funding transaction succeeded, but hasn't yet been confirmed by the ledger the deploy step queries. main.rs's own comment on the retry constants names this explicitly as one of the two reasons deploy retry exists:

"when friendbot funding is suspected to have failed transiently (rate-limiting, network hiccups, or the account not being fully confirmed on-ledger yet)"

(from source, quoted directly.)

Cause: Ledger propagation delay between "friendbot accepted the funding request" and "the account is visible to the network for the deploy transaction."

What to do: Identical to rate limiting — this is exactly what the exponential backoff is for. No action needed beyond letting the retries run; a source identity that was already funded before this run won't hit this path at all.

Friendbot / testnet unavailable or unreachable ​

Symptom: Deploy fails (immediately or after retries) with guidance that is not the rate-limit one. The final classifier separates two sub-cases by the error text:

Cause: the deploy/funding service returned a server error — it is down or
overloaded, not something on your side.
  What to do: check https://status.stellar.org, then re-run in a few minutes.
Cause: the network could not be reached (DNS / connection / timeout).
  What to do: check your own connectivity, any proxy or VPN, and firewall rules
  for outbound HTTPS to Stellar infrastructure, then re-run.

Cause: The first is a 5xx / "unavailable" response — testnet infrastructure is down or overloaded. The second is a connection-level failure (DNS, refused connection, timeout) — the request never reached a server, which more often points at your own egress than at testnet.

What to do: For the service-error case, check Stellar's status page and wait it out — no flag change helps ride out an outage. For the unreachable case, the problem is between you and the network: check connectivity, proxy/VPN, and outbound-HTTPS firewall rules.

RPC unreachable (simulateTransaction) ​

Symptom: Simulation fails — either after retries, or immediately depending on the specific network condition — with a message shaped like:

simulateTransaction RPC failed: curl exited with status <code>: <stderr>

(from source — the exact wrapper text in LiveTransport::simulate_transaction, live.rs.)

Cause and reproduction: This tool shells out to curl -s -X POST ... https://soroban-testnet.stellar.org:443 directly (see the --network discrepancy note — this URL is hardcoded and does not follow --network). Reproduced directly in this environment: curl's -s/--silent flag suppresses its progress meter and its own error text, so a DNS or connection failure surfaces here as an exit status with an empty <stderr> — not silence from the tool hiding something, but curl genuinely not printing anything under -s:

$ curl -X POST -H "Content-Type: application/json" -d '{}' https://this-host-does-not-exist.invalid:443
curl: (6) Could not resolve host: this-host-does-not-exist.invalid
$ echo $?
6

With -s (as this tool invokes it), that run produces exit code 6 and empty stderr — so the message you'll actually see is closer to curl exited with status: 6 with nothing after the colon. The exit code is still the fastest way to identify the cause; the common ones from curl(1):

Exit codeMeaning
6Could not resolve host — DNS failure, no route to the RPC host
7Failed to connect — host resolved but refused the connection, or a firewall/proxy is blocking it
28Operation timeout — the host is reachable but not responding
35SSL/TLS connect error

What to do: This class of failure is in the retry classifier's transient list (connection, timed out, timeout, unavailable all match curl's own wording for these cases), so it retries automatically. If it still fails after all attempts: confirm outbound network access to soroban-testnet.stellar.org:443 from wherever the tool is running (a CI runner's egress firewall is a common culprit — this is exactly the situation this project's own budget.yml sidesteps by mocking the Tier B step for fork pull requests, since forks don't have the network access or secrets Tier B needs), and confirm curl itself is on PATH and working (curl --version).

Simulation failure (transaction simulation failed or similar) ​

Symptom: The run completes for other functions, but one specific function is skipped with a warning and reported as failed rather than crashing the whole run. The exact warning line depends on which of three sub-cases occurred (from source, the match on SimulationFailure around line 1684 of main.rs):

Warning: Simulation failed for <function>: <stellar CLI stderr>
Warning: RPC error for <function>: <RPC "error" field contents>
Warning: Failed to extract metrics for <function>: <parse error>

The report still prints for every function that succeeded; a fully failed run instead says No successful simulations to report. and exits 0 (from source — main.rs's handling around SimulationOutcome::Failed and the "no successful simulations" message).

Cause: simulate_function classifies this into the same three sub-cases as the warnings above, all reported as SimulationFailure rather than aborting the process (from source, error.rs):

  • Invoke (Warning: Simulation failed for ...) — stellar contract invoke --build-only itself failed (bad arguments in [functions.&lt;name&gt;].args, a function name that doesn't match the deployed contract's actual signature, etc.).
  • Rpc (Warning: RPC error for ...) — the simulateTransaction response body parsed fine, but its JSON carried an "error" field — the network answered, and the answer was "no." This is the case the --network discrepancy produces when you deploy to a non-testnet network: the simulate step queries testnet for a contract ID that only exists on the network you actually deployed to, and testnet correctly reports it as unknown.
  • MetricsExtraction (Warning: Failed to extract metrics for ...) — the RPC responded successfully, but the expected SorobanTransactionData fields weren't where the tool expected them (a Protocol/SDK mismatch is the likely cause — see the supported-versions table).

What to do: Check which sub-case you're in from the surrounding warning text. Invoke failures are almost always a budget.toml [functions.&lt;name&gt;].args problem — cross-check against the function's real signature. Rpc failures where you're confident the contract and function are right are the --network trap above — confirm you're on testnet. MetricsExtraction is worth filing as an issue; it usually means an SDK/protocol version this tool hasn't been updated for.

A contract that exports nothing simulatable ​

The tool discovers what to simulate by parsing each contract's compiled WASM export section. When that yields nothing usable there are three distinct causes, and each now produces its own message naming the package and saying what to change (from source, the contract_exports module and its call sites in main.rs / watch.rs).

Cause 1 — the crate isn't a cdylib. A crate that pulls in soroban-sdk as a normal dependency but whose [lib] crate-type doesn't include cdylib produces no WASM at all. It is still skipped (there is nothing to build), but no longer in silence:

Package '<name>' depends on soroban-sdk but its `[lib] crate-type` does not
include `cdylib`, so it produces no WASM and cannot be measured. Skipping.
  Add `crate-type = ["cdylib"]` to its `[lib]` section (keep `rlib` too if
  other crates depend on it) if it is meant to be a contract.

Fix it in Cargo.toml:

toml
[lib]
crate-type = ["cdylib", "rlib"]

Cause 2 — a cdylib whose WASM has no function exports at all. The crate compiled as a plain library, or the #[contract] / #[contractimpl] macros were never applied:

Error: Package '<name>' built a WASM with no function exports.
  The crate most likely compiled as a plain library, or the Soroban contract
  macros (#[contract] / #[contractimpl]) were never applied.
  Put #[contractimpl] on the contract's impl block and confirm its `[lib]
  crate-type` includes `cdylib`, then rebuild.

Cause 3 — a cdylib that exports functions, but none are contract entrypoints. Every export is a toolchain symbol (_start, __data_end, …). The message lists what it found, so you can see the mismatch:

Error: Package '<name>' exports functions, but none match the Soroban contract
calling convention.
  Exports found: _start
  Those are toolchain symbols, not contract entrypoints. Add #[contractimpl] to
  the contract's impl block so its methods are exported under their own names,
  then rebuild.

Causes 2 and 3 are treated as run failures: a crate deliberately built as a cdylib that produces no contract entrypoint is a real misconfiguration, so the run exits non-zero rather than reporting "no successful simulations" and exiting 0.

Missing or misconfigured identity ​

Symptom: Deploy fails on the very first attempt — no retry lines — with a stellar contract deploy failed: ... message (from source, LiveTransport::deploy_contract's error wrapper) whose actual text comes from the stellar CLI itself, not this tool. The tool cannot predict that text since it depends on the installed Stellar CLI version and your local identity configuration — this environment does not have the stellar CLI installed to reproduce it directly, so treat the CLI's own wording as authoritative over any paraphrase here.

Cause: The identity named by --source (or source in budget.toml) either doesn't exist in your local Stellar CLI's keystore, or exists but isn't funded on the target network.

When the CLI's error text carries a recognisable "underfunded / account not found" signal, the final classifier picks it out and prints the account-specific guidance, substituting your actual --source and --network:

Cause: the source account 'alice' is missing or unfunded on testnet.
  What to do: fund it and re-run — `stellar keys fund alice --network testnet`
  (or `stellar keys generate alice --network testnet --fund` to create it). This
  will not resolve by waiting.

What to do:

  • Confirm the identity exists: stellar keys ls.
  • Confirm it's funded on the network you're using: stellar keys fund &lt;name&gt; --network testnet (or generate + fund in one step: stellar keys generate &lt;name&gt; --network testnet --fund, as the End-User Guide shows).
  • Because this fails on attempt 1 with no retry, it will not self-resolve by waiting — unlike the friendbot cases above, this needs a configuration fix before re-running.

Stellar CLI or the wasm32 target isn't installed ​

Not a network failure at all, but the tool checks for both before doing any network work (run_preflight_checks in main.rs, skipped entirely under --replay), specifically so a missing local tool doesn't masquerade as a network problem:

Stellar CLI is not installed or not on PATH.
Install it with:  cargo install --locked stellar-cli
See: https://github.com/stellar/stellar-cli
wasm32-unknown-unknown target is not installed.
Install it with:  rustup target add wasm32-unknown-unknown

(from source, both exact strings.) These fail immediately with no retry, since no amount of waiting installs a binary.

Summary: which page to reach for ​

  • Wrong output, right network (a check failed, a limit needs updating): see the main Tool Reference, not this page.
  • Nothing works and you don't know why: work top-to-bottom through First: is it hung, or is it retrying? to classify the failure, then jump to the matching entry in the Failure catalog above.
  • Every flag's exact behavior: the CLI flag reference documents all of them, including two behaviors (--color, --network) that this page and that one both link to because they directly affect how testnet failures show up.

Built for the Stellar & Soroban ecosystem.