Skip to content

recourse

recourse (n.): a source of help or strength.

recourse coordinates resilience for Go services. Retries, hedging, circuit breaking, timeouts, and budgets work together under a shared policy, with structured visibility into execution.

Call sites provide a stable policy key. Local or remote policies define how calls handle failures, slow responses, and additional load. Teams can tune that behavior consistently across services without rebuilding resilience controls at every call site.

New here? Start with Design overview, then move to Getting started.

Why recourse?

A slow or failing dependency forces several decisions: whether to wait, start another attempt, reject the call, or limit additional work. Those decisions affect both latency and load.

recourse brings these controls into one execution path:

What you need to control What recourse provides
Transient failures Retries with bounded attempts, backoff, and classifiers that decide which outcomes warrant another attempt.
Slow responses Hedging that can start additional attempts after a fixed delay or a latency-based trigger.
Repeated dependency failures Circuit breaking that rejects calls while the circuit is open and probes for recovery after a cooldown.
Additional load Budgets that allow or deny retry and hedge attempts.
Time spent on a call Per-attempt and overall timeouts configured through policy.
Incident diagnosis Timelines and observer hooks that expose attempt outcomes, timing, and execution decisions.

The controls also account for one another. When a circuit breaker probes for recovery, hedging is disabled so the probe does not launch extra requests. Configured budgets limit the additional work from retries and hedges.

What “policy-driven” means

Call sites supply a stable, low-cardinality key, such as "payments.Charge". The provider resolves a policy for that key, and the executor applies its configured behavior:

  • Attempt limits, backoff, and jitter.
  • Hedge limits and triggers for starting concurrent attempts.
  • Circuit breaker thresholds and recovery cooldowns.
  • Per-attempt and overall timeouts.
  • Classifier selection to interpret errors and results.
  • Budgets to gate retries and hedges.

Policies can live in your application through StaticProvider or come from an external source through RemoteProvider. The remote provider caches policies, coalesces concurrent fetches, and offers opt-in last-known-good fallback during source failures. See Remote Configuration.

This separates the operation from its resilience configuration: the call site keeps doing its work while the policy defines how that work is protected.

Quick start

The facade API takes a string key like "user-service.GetUser":

package main

import (
    "context"

    "github.com/aponysus/recourse/recourse"
)

type User struct{ ID string }

func main() {
    user, err := recourse.DoValue[User](context.Background(), "user-service.GetUser", func(ctx context.Context) (User, error) {
        // call dependency here
        return User{ID: "123"}, nil
    })
    _ = user
    _ = err
}

This minimal example uses the default policy, which provides bounded retries. Hedging, circuit breaking, timeouts, and budget limits are configured explicitly. See Getting started to install your policies and executor at startup.

When you need to know what happened, request a timeline:

ctx, capture := observe.RecordTimeline(ctx)
user, err := recourse.DoValue(ctx, "user-service.GetUser", op)
_ = user
_ = err

tl := capture.Timeline()
for _, a := range tl.Attempts {
    // a.Attempt, a.Outcome, a.BudgetAllowed, a.Backoff, a.Err, ...
}

Observability-first

When a call fails or takes longer than expected, you need to know which policy ran, which attempts were launched, and what stopped further work.

recourse captures a structured observe.Timeline (attempt timings, outcomes, budget decisions, errors) and can also stream attempt/timeline events to your own logging/metrics/tracing via observe.Observer.

What’s inside

  • Policy keys: stable, low-cardinality keys ("svc.Method") that select behavior.
  • Policies + providers: policy.EffectivePolicy resolved through in-process controlplane.StaticProvider or cached controlplane.RemoteProvider lookups.
  • Execution: retry.Executor coordinates retries, hedging, circuit breaking, and timeouts according to policy.
  • Classifiers: pluggable (value, err) → Outcome decisions to retry, stop, or abort based on protocol and domain semantics.
  • Budgets/backpressure: gates on retry and hedge attempts to control additional load.
  • Observability: structured observe.Timeline plus streaming hooks via observe.Observer.

Where to go next