PUNKthe adaptive runtime

//DOCS What's New

User-facing record of what shipped in Punk, newest first.

What's New

Customer-facing monthly highlights for Punk. We group related work and include only changes that materially affect adoption, compatibility, reliability, governance, or measured value. Engineering-level fixes are intentionally omitted.

August 2026

Faster First-Run Activation

Generated starter smoke tests now treat run-detail lookup as advisory after both answers, run ids, route evidence, task correctness, and exact replay have been proven. A temporary detail-projection outage produces explicit ordered warnings instead of a false activation failure.

Generated starters now verify that both hosted smoke-test answers actually satisfy the billing-classification task before recording success. Harmless casing, whitespace, punctuation, code fences, and emphasis remain accepted, while wrong labels, explanations, inconsistent exact repeats, and missing route proof fail machine-readably.

Tenants created after a worker starts now activate newly scheduled workflows and agents immediately across create, edit, and import paths, without waiting for a restart or unrelated scheduler cycle.

Dashboard user provisioning now uses semantic email, name, role, and temporary-password controls. Temporary passwords stay masked and bounded, password managers receive the correct new-user hints, and the form stacks without horizontal overflow on narrow screens.

The interactive offline demo now leaves its locally started gateway and populated dashboard available after the savings report, then shuts it down cleanly on Ctrl-C. Non-interactive runs still clean up automatically, and a demo never retains or stops a gateway it did not start.

Malformed bootstrap administrator identities and weak bootstrap passwords now stop startup before Punk creates an organization or user, while valid local-domain and internationalized addresses remain supported. Home also keeps core run and savings evidence available when optional opportunity or learning-history reads are degraded.

Broader Compatibility

OpenAI Chat Completions usage now retains documented audio and accepted or rejected prediction token details across live, streaming, and exact-cache responses. Contradictory token aggregates fail before optimized reuse instead of becoming durable billing evidence.

Deterministic OpenAI Responses requests can now reuse exact results within one safety_identifier while different identities stay isolated. Punk preserves opaque safety identifiers byte-for-byte for the provider instead of imposing a gateway-specific length or whitespace contract.

OpenAI streaming callers can now send stream_options.include_obfuscation: false through Punk unchanged. Requests that require provider-added padding still fail before provider work because Punk reserializes SSE and cannot preserve that wire contract.

Required Anthropic service tiers and context management now stay attached to the Anthropic provider during recovery, while documented prompt-cache hints and the default auto tier retain cross-provider fail-open availability in buffered and streamed calls.

Anthropic citation-bearing requests and active OpenAI provider extensions now remain attached to their provider family during recovery, preventing provenance or extension semantics from being silently discarded. Stateless Responses requests, neutral controls, and documented prompt-cache hints still fail over across healthy provider families.

OpenAI Chat Completions requests can now constrain automatic or required selection to a named subset of declared tools. Punk forwards the documented allowed_tools contract byte-for-byte on live requests and rejects any provider response that attempts a tool outside the permitted subset.

Provider-owned continuation state now remains attached to its provider family: non-null OpenAI Responses previous_response_id requests and non-null Anthropic containers fail explicitly when that family is unavailable, while stateless requests retain healthy cross-provider fail-open behavior.

OpenAI assistant audio references now continue through Punk as provider-bound live requests. Opaque provider ids pass through byte-for-byte, contribute only a redacted hash to trace identity, and never enter response caches, artifacts, substitution, or cross-provider failover.

More Trustworthy Runtime Evidence

Workflow runs now snapshot their redaction controls once and apply that decision consistently across trace and governance evidence. Successive-input learning reads trace histories in tenant/app-scoped batches, and pattern inventory facets use one grouped snapshot, preserving measured optimization quality and isolation while reducing storage-bound latency.

Abandoned OpenAI and Anthropic response streams now cancel their established upstream provider work when the downstream reader disconnects. Partial streams finalize as failed and cannot seed cache, evaluation, or learning evidence.

Malformed legacy identity values are now concealed across signed-in account, organization, and audit evidence without rewriting the append-only ledger. Completed live answers also stay available when a failed artifact or learned-route accounting write stalls, and large run inventories compute complete tenant/app-scoped facets in one aggregate read.

API-key revocation now commits together with its audit receipt and reports confirmed, unchanged, or outcome-unknown results explicitly under failures and concurrent requests. Governance screens also conceal legacy identity values that fail current validation—including the signed-in footer—while preserving ID-bound repair and removal actions. Buzz operator evidence loads recent outbox state in one tenant-scoped batch instead of one request per event.

Tripwire inventory and qualified model-substitution cadence reads now share and cap stalled storage work instead of pinning customer traffic indefinitely. Slightly delayed valid evidence still preserves governance blocks and cheaper-model routes; a true outage falls back with explicit route evidence and recovery starts from a fresh read.

App-pinned administrative credentials can no longer enter tenant-wide user and organization identity controls. Tenant-wide administrators and signed-in organization operators retain their existing member, rename, and invitation workflows.

External prompt ingestion and savings evidence now stay responsive and tenant-isolated when storage stalls. Late definitive ingest failures receive one safe idempotent retry, slightly slow valid savings remain available, and Home and Results preserve current request and recovery evidence while naming a savings outage explicitly.

Exact-cache route discovery now reserves retained capacity per tenant during storage stalls, so one noisy tenant cannot turn another tenant's proven repeat into a paid call. Concurrent learning-report reads also share one bounded projection request and retain the latest valid evidence when storage is delayed or rejects.

Route discovery and usage projection now stay bounded and tenant-isolated during storage stalls. A noisy tenant cannot consume another tenant's retained discovery capacity, paid live answers preserve their authoritative status and body, and incomplete usage evidence never enters optimized reuse.

Completed dashboard chat replies now remain available when advisory run summaries stall, and one tenant's outage cannot consume the process-wide projection budget or erase another tenant's real route and cost evidence. Cache-audit retention is likewise partitioned so healthy tenants keep their own cache evidence during a noisy tenant's storage outage.

The newest operator review of an artifact-served run is now the only feedback signal allowed to change artifact confidence and lifecycle evidence, even when an older request resumes through another storage owner after its trace acknowledgement stalls.

Retried or overlapping model-substitution shadows now contribute one immutable quality sample and one route statistic. Anthropic prompt-cache writes and reads are reflected at their TTL-specific list prices in stored cost and savings evidence, including compatibility with nullable provider counters.

Paid live answers now finish within a bounded evidence window when best-execution or post-live shadow storage stalls. Punk marks incomplete or indeterminate evidence explicitly and prevents it from entering cache, shadow, or learning reuse until proof is complete.

Buffered OpenAI Chat Completions can now return requested token log probabilities through Punk's live path; mismatched token-byte evidence is rejected and DLP changes suppress stale confidence metadata. Paid buffered answers also remain available when completion evidence storage stalls, with tenant-scoped retained-work limits, while cache hits require the complete normalized wire snapshot to remain unchanged before serving.

Provider readiness now distinguishes configured credentials from recent live gateway proof, ignores imported or cross-tenant evidence, and gives operators a concrete recovery action after stale or failed traffic. Provider-key replacements preserve the working credential on failure and commit rotations with their exact audit evidence; password changes invalidate every other browser session.

Delayed artifact inventory and preference-evidence reads fail open within a fixed bound, cap retained work across repeated bursts, and revalidate lifecycle state so a stale promoted snapshot cannot route after quarantine.

Authenticated traffic and governance denials now remain bounded through credential-index, memory-influence, and tripwire-evidence outages. Unavailable key authority fails retryably closed without provider calls, memory quarantine retains approval-required behavior with capped fallback work, and detected tripwires stay hard-blocked while advisory event writes settle.

Home remains usable when background-job telemetry is unavailable, usage reports attribute complete requested windows without double-counting concurrent activity, and cache hits or live fallback stay responsive during delayed audit writes while recovered evidence still settles durably.

Large run evidence packets now download completely even when they exceed the dashboard's interactive response limit. Cache routes report savings and avoided latency from the observed source call—including legitimate zero-cost and zero-latency evidence—instead of substituting estimates. Provider-affine OpenAI requests now fail explicitly when their required provider family is unavailable, rather than returning a simulated mock answer; ordinary non-affine requests still retain healthy live-provider failover.

Complete SDK evidence exports now stream with backpressure, cancellation, and independent response limits, so slow consumers can retrieve large packets without unbounded buffering. Optimization proof now includes sampling controls; existing proof safely falls back live once and re-earns evidence instead of serving an answer produced under different sampling behavior.

Gateway failures now include browser-readable request ids that typed SDK errors preserve for support correlation. Anthropic artifact routes conservatively honor output caps and fail open to the live provider, scheduled worker ticks reject unknown or ambiguous controls before starting work, and OpenAI SDK streams request usage while retaining complete terminal model, provider, and token evidence across live and optimized routes.

Semantic-answer reuse now partitions deterministic values across every conversation role, Anthropic SDK streams preserve monotonic token and cache-accounting evidence, positive SDK notes cannot become corrective failures, and usage reports reject ambiguous windows instead of silently changing them.

July 2026

A Faster Path To Measured Value

  • The Overview now leads with what needs attention, the next recommended action, and progress from first traffic to a proven optimization, then asks operators for the safe repeat that demonstrates reuse. First-time operators also get an app-specific connection receipt that distinguishes received traffic from a served run without exposing sibling applications; SDK applications can submit typed named metrics against settled runs.
  • Trial, onboarding, agent activation, and design-partner paths are clearer: the offline demo starts and cleans up an isolated mock gateway and waits for its complete shadow batch, generated starters preflight file collisions and verify fixed-task answers before recording positive feedback, hosted smoke tests separate required route proof from optional savings, expose observe-mode would-have savings without claiming they were served, reject inconsistent exact replay, and keep onboarding evidence shareable. Codex diagnostics now distinguish a broken installation from an unsupported version and provide a direct recovery action, while the offline demo names the managed authority required for production tool-result reuse.

Proof That Is Easier To Inspect

  • Run, pattern, and artifact views explain routing decisions in plain language while keeping replay, shadow, policy, and fallback evidence available for deeper review; app-scoped run evidence is filtered before bounded result windows, concurrent replay requests converge on one immutable proof, counterfactual audits cannot mix another tenant's evaluations, and ledger proof distinguishes complete hash chains from compatible partial legacy coverage. SDK clients now expose and validate complete route explanations and memory-influence acknowledgements against the requested run, retain server-owned tool-cache ineligibility evidence, and can correlate exact LangChain response headers through a failure-contained onRun observer.
  • Punk published a measured public reference study covering 150 benchmark requests and 50 held-out treatment requests. It is explicitly presented as mechanism proof—not customer or production evidence.

Safer Production Behavior

Concurrent administrators can no longer remove every owner from an organization, even when requests run through separate storage instances. Production readiness checks now settle with explicit dependency timeout evidence and share stalled probes across polling bursts instead of retaining one storage read per poll.

Concurrent workflow and agent creation now reserves plan capacity atomically across runtime instances, while validated API-key and browser-session traffic stays available through stalled activity bookkeeping without retaining one write per request.

  • Cache invalidation now commits with its operator receipt or rolls back, including recovery from a lost storage acknowledgement without purging newly created entries. Legacy clients can request fresh work with Pragma: no-cache when modern cache controls are absent, and MCP registry mutations reject misspelled or server-owned fields before applying configuration. Invalid Stripe, Chargebee, and Shopify financial-action caps fail closed before level-4 plans; concurrent dashboard turns retain complete alternating history; concurrent replay requests commit one proof; abandoned inbound bodies settle when callers disconnect; and recovered Anthropic /v1/v1 paths preserve canonical validation-first budgets. Shopify, Twilio, and Mailchimp reads now stop stalled or oversized upstream responses while preserving successful JSON that arrives without a media-type header; accepted oversized connector writes retain request identity and are explicitly marked outcome-unknown and unsafe to retry. Optimized routes fall back to the configured live provider more consistently when evidence is missing, stale, inconsistent, or unsafe; healthy provider backups remain reachable through rejected or stalled failover-audit writes; hard MeshGuard denials override approval hints, observe-mode SDK writes execute with audited evidence while optimize-mode gates remain enforced, provider refusal/content-filter terminals remain successful but non-reusable across buffered and streamed protocols, throttled chat requests retain blocked-run explanations, undeliverable team invitations are revoked with their denial audit, live answers remain available during request, terminal, model-completion, or run-projection evidence outages without seeding future optimization, provider-requested storage stays live and bound to the selected provider, validated keys and sessions survive activity-timestamp outages, canary promotions pause when rollout policy state cannot be read, and opportunity estimates use retained 30-day live work rather than stale lifetime bursts.
  • Abandoned dashboard chats and manual workflow or agent runs now stop downstream provider work instead of continuing after the caller leaves; onboarding distinguishes served traffic from side-loaded history, and large run timelines stay within the dashboard response budget while retaining complete export evidence. Tenant settings and opened web sessions now commit with exact audit proof or roll back; cancelled MCP writes retain secret-safe outcome-unknown/no-retry evidence; and durable webhook retries expose stable bounded keys for receiver deduplication. Transient run-start writes now recover without duplicate paid work, while rejected or stalled model-request and DLP audit writes stay bounded without blocking healthy providers or truncating already-masked answers. Negative reviews now remove the affected artifact from routing immediately, including after a retried lifecycle fault, ambiguous workflow names require an explicit workflow id instead of silently choosing a target, incomplete provider outputs stay out of learning and replay proof, artifact evidence honors trace redaction without changing served answers, and Chorus output tripwires inspect assistant text plus tool arguments while retaining honest incurred-cost evidence. SDK side-effect traces retain idempotency evidence through terminal projections without hiding distinct tool executions, and Linear calls stop stalled or oversized responses while keeping retries read-safe and timed-out writes outcome-unknown.

  • App and tenant boundaries, approval flows, audit reasons, permission-fingerprinted tool caches, and high-impact lifecycle controls are more conservative and easier to verify; app-pinned administrators now see only their app's agent identities and live web sessions, see and decide only their app's approvals, and cannot inspect or mutate tenant-wide settings, credentials, connector setup, MCP-server configuration, or tenant-wide tripwire controls, individual API keys can move atomically between optimize and observe modes, and administrators can recover passwords through enumeration-neutral reset links that invalidate prior sessions.
  • Runtime telemetry, SLO checks, load testing, bounded provider, MeshGuard, connected MCP, Slack/shared connector, SDK, SEC research, and GitHub waits and response bodies, private SDK no-store requests, post-header cancellation, usage-accounting fail-open continuity, atomic recent-window dashboard chat turns, outcome-unknown non-retryable write timeouts, write-aware scheduled retries, redirect-contained workflow notifications, retryable MCP OAuth recovery, and dead-lettered crashed leases make production readiness and operational risk clearer; non-cooperative workflow writes cannot overrun their deadline, bounded scheduler stalls no longer duplicate live jobs, MCP explanations retain readable rejected-route evidence, Stripe setup contracts fail before success, retention removes run evidence atomically, confirmed email delivery survives stalled response cleanup, and the dedicated Buzz worker verifies an already-migrated schema without DDL and starts commissioning before its live evidence becomes the release-readiness gate.

Broader Compatibility

Anthropic citation-enabled document requests now retain document, web-search, and custom search-result citations through buffered and streamed live responses. Citation metadata is bounded, covered by DLP and tripwire enforcement, accepts nullable provider titles and stable non-URL retrieval identifiers, and stays outside response caches.

OpenAI streaming responses now preserve requested token log probabilities with response-wide bounds and correct reconstruction when one Unicode character spans adjacent token-byte arrays. Copied base URLs that already end in /v1/chat/completions also recover without producing a duplicated /v1 route.

Buffered OpenAI responses now preserve validated URL citation annotations and exact Unicode spans on the live path. Citation metadata is covered by DLP and tripwire enforcement and stays outside response caches so provider provenance cannot be replayed as local evidence.

Current OpenAI and Anthropic browser SDKs can call Punk directly with their official metadata headers, and SDK run inspection now exposes the latest byte-bounded event window instead of failing once a valid trace exceeds the client's response ceiling.

  • The integration catalog now includes 51 managed SaaS connectors and 517 governed tools across engineering, support, data, finance, security, and business systems.
  • Structured OpenAI output contracts are checked before buffered or streamed answers are served, generated tool schemas can use bounded local JSON Pointer references (including URI-encoded fragments), validated json_object requests can use promoted artifacts without a provider call while non-object results fail open, and Google Workspace calls now use bounded deadlines while large document exports retain a safe truncated prefix.

  • MCP setup now accepts copied OpenAI and Anthropic endpoint URLs while preserving nested reverse-proxy roots, and the Runs view can filter exact serving-provider evidence across scoped pagination. OpenAI and Anthropic compatibility now includes native Claude image blocks with canonical fallbacks for unsupported references, native tool-call streaming, safe cross-protocol streamed-plan reuse, stop-control preservation in both cross-provider directions, matched Anthropic stop-sequence evidence across live, cached, and streamed responses, serial tool-call intent across buffered and streamed failover, cache isolation for distinct OpenAI token-limit controls and ordered tool catalogs, live-sibling failover before mock fallback, cancellable and output-bounded provider streams with prompt upstream release, runtime rejection of invalid declared text without breaking empty tool-only responses, independent bounds for oversized provider bodies and non-cooperative polling transports, surfaced completion/stop reasons, complete bounded streamed tool-call assembly, buffered and streamed refusal fidelity with explicit SDK safety-stop evidence, signed Claude thinking continuity across reusable tool plans without learner exposure, served-model fidelity across live/failover/cache paths, deterministic repeat-cache defaults in generated Anthropic starters, preserved assistant refusal history and provider backoff hints, bounded stalled-stream recovery, streamed OpenAI usage/completion proof, streamed reuse of eligible buffered exact responses, and standard Cache-Control: no-cache, no-store, and zero-age revalidation controls for fresh live work.
  • Standard OpenAI completion metadata now passes through byte-faithfully, follows provider character limits, and keeps those requests on live routes so cache reuse cannot erase provider-visible metadata. Repeated OpenAI streaming requests can now receive eligible semantic-cache answers as valid SSE without another model call, current OpenAI safety identifiers isolate optimized reuse, and requests using different Anthropic API versions no longer share cached response bytes.

  • OpenAI, Anthropic, OpenRouter, Vercel AI SDK, LangChain, and Claude Code adoption paths now have clearer setup guidance, safer launch validation, examples, and compact SDK route proof with fallback for older gateways; TypeScript SDK tool continuations and multimodal text/image/audio/file inputs compile without casts, cross-provider top_p stays intact, and native nullable container, context-management, and service-tier response state survive live, exact-cache, and synthesized Anthropic streams while stateful containers remain live-only.

Easier Daily Operation

  • A keyboard command palette, better linked evidence, clearer filters, prefix-scoped cache eviction, safer lifecycle actions, SDK rollback/quarantine and trustworthy thumbs feedback, and resumable ignored patterns reduce the effort required to navigate and operate the runtime.

June 2026

Adaptive Runtime Foundation

  • Punk launched with OpenAI-compatible and Anthropic-compatible gateway surfaces, policy enforcement, route explanations, provider failover, trace evidence, and cost tracking.
  • Repeated work can move through evidence-gated cache, model substitution, and deterministic artifact routes, backed by replay, shadow, canary, rollback, quarantine, and approval controls.

Operating Surface

  • The initial dashboard brought together runs, patterns, artifacts, learning, savings, governance, workflows, agents, and approvals, with hosted organizations, API keys, usage metering, and billing support.