What's New
Customer-facing monthly highlights for Punk. We group related work and include only changes that materially affect adoption, compatibility, reliability, governance, or measured value. Engineering-level fixes are intentionally omitted.
September 2026
Demo setup now repairs only an untouched workflow reservation left by an interrupted first run. Disabled workflows and administrator edits remain unchanged, while equivalent older reservations still recover when their stored JSON uses a different key order.
Workflow and agent lists now apply their kind filter before the storage page limit. Older matching rows no longer disappear behind newer rows of the other kind.
Artifact evaluations now show the stored evaluation identity. Missing or unusable ids stay unavailable instead of inventing a proof record from the compared run or artifact.
Artifact details now show when proof was last evaluated. Missing or non-numeric timestamps stay unavailable instead of inventing recency from the created time.
Artifact details now show the stored estimated serving cost and latency. Missing or non-numeric estimates stay unavailable instead of inventing a cheaper or faster path.
The aggregate self-serve activation report now shows median time from first successful gateway completion to the earliest qualifying distinct second completion. Concurrent requests use the first success operators actually receive, and impossible values beyond the 24-hour cohort remain unavailable.
Canary artifact details now show the live and shadow counts since the current rollout percent began. If the baseline captured at that percent is missing, or current counters are behind it, the page says the rung window cannot be measured instead of inventing progress.
Typed SDK artifact evaluations now preserve one validated record across artifact detail and replay. Inherited fields, accessors, and later prototype changes cannot forge or change evaluation evidence.
Artifact evaluations now show the compared run. Missing or unusable run ids stay unavailable instead of inventing a trace link.
Artifact evaluations now show the stored side-effect pass or fail. Missing or non-boolean side-effect evidence stays unavailable instead of looking like a pass.
Typed SDK pattern and artifact records now come from complete own-data snapshots. Inherited fields, accessors, and later prototype changes cannot forge optimization identity, state, evidence counters, or pattern links across list, detail, replay, and lifecycle responses.
OpenAI Responses served through an adapted provider now preserve admitted metadata across live and exact-cache routes. Gateway release-check retries use a fresh measured identity and must prove live plus reused coverage again, so a warm cache cannot hide a slow live path.
The aggregate self-serve activation report now measures median time from creation of the exact Punk API key to its first successful gateway run. The privacy-safe measure comes from append-only run evidence, so removing an operational key cannot erase the customer journey.
Managed connector tool discovery now applies its installation limit in storage and reports when more installations exist. Empty connector filters keep their documented unfiltered behavior.
OpenRouter runs now use the exact charge returned by OpenRouter for buffered and streamed requests. When upstream charge evidence is unavailable, Punk continues to use its bounded local estimate.
Connect and Home setup checks now distinguish provider-backed hybrid and model-substitution routes from deterministic shortcuts. Chorus remains explicitly unknown until its receipt proves which execution path ran.
August 2026
Faster First-Run Activation
The Chorus OpenAI SDK example now preserves each server-owned dashboard run URL. Split gateway and dashboard deployments open the right run, while missing, mismatched, credential-bearing, query-bearing, invalid, and oversized receipts use a safe gateway fallback.
Connect now recognizes when a second request already matched an observed plan cache. It directs operators to review that evidence before optimize mode instead of repeating the generic replay-and-shadow instruction, while live, blocked, mismatched, and missing evidence remain conservative.
Generated starters and punk smoke now preserve validated dashboard run links from confirmed successes and failures across canonical and equivalent encoded routes. Unconfirmed trace-admission failures keep correlation and retry evidence without advertising an unrecorded run, while unsafe or mismatched receipts still use the safe gateway route.
Copied Home and Connect setup checks now keep the dashboard destination returned by Punk for both new runs. Split gateway and dashboard deployments open the right host, while older or unsafe receipts use the existing gateway link.
Direct gateway responses now include the exact dashboard run URL beside the Punk run id and route across buffered and streamed OpenAI Chat, OpenAI Responses, and Anthropic calls. Dashboard Chat and server-owned workflow diagnostics remain visible in traces but no longer count as connected-agent activation.
Customer-escalation handoffs now include validated HubSpot ARR in Zendesk and Slack. Missing or unsafe values remain explicit as Not available, and existing saved workflows upgrade without changing customer-owned graphs.
Self-serve 24-hour repeat activation now requires the second distinct input to come from the same app and agent as the first successful run. A different agent in one app no longer inflates repeated-workflow activation, while legacy runs with no agent identity match only other identity-free runs.
punk smoke now reports the second distinct request's route, whether the response still used a provider model, what proof is still required, and the matching dashboard run URL. Provider-backed optimized routes remain explicit, while blocked or distributed routes use conservative unknown evidence.
punk smoke --stream now verifies streamed OpenAI Chat Completions, OpenAI Responses, Anthropic Messages, and OpenRouter paths. It requires a successful terminal, stable run and route evidence, a fresh first request, and a byte-faithful exact-cache repeat, so a deployment with broken streaming no longer passes activation smoke.
The public GitHub Actions auto-triage pilot now includes a repeatable local experiment, governed dry-run workflow, deterministic artifact, and checked-in evidence. On the 40-issue public fixture, optimize mode preserves exact labels while cutting model calls by 50%, model cost by 14.82%, and modeled latency by 22.62%; the result is simulation evidence, not a production deployment claim.
Results now surfaces workflows whose repeatable value is lower waiting time even when model pricing is zero. These opportunities are labeled as estimates, remain separate from recorded savings, and preserve the existing cost-first ranking when the view is full.
Self-serve activation now recognizes a completed optimized run with verified latency savings even when model pricing produces no dollar savings. Cost-only, latency-only, and mixed workflows therefore share one honest time-to-first-value measure.
Generated Responses starters now stop declared or streamed success bodies above 8 MiB before unbounded parsing, cancel unread bytes best-effort, and still accept valid JSON at the exact UTF-8 ceiling.
Generated starters now support native OpenAI Responses for both smoke and agent entrypoints with the currently published SDK. First and repeat requests retain distinct run ids and route evidence, including exact-cache proof.
Generated starter smoke tests now treat run-detail lookup as advisory after both answers, run ids, route evidence, task correctness, and exact replay have been proven. A temporary detail-projection outage produces explicit ordered warnings instead of a false activation failure.
Generated starters now verify that both hosted smoke-test answers actually satisfy the billing-classification task before recording success. Harmless casing, whitespace, punctuation, code fences, and emphasis remain accepted, while wrong labels, explanations, inconsistent exact repeats, and missing route proof fail machine-readably.
Tenants created after a worker starts now activate newly scheduled workflows and agents immediately across create, edit, and import paths, without waiting for a restart or unrelated scheduler cycle.
Dashboard user provisioning now uses semantic email, name, role, and temporary-password controls. Temporary passwords stay masked and bounded, password managers receive the correct new-user hints, and the form stacks without horizontal overflow on narrow screens.
The interactive offline demo now leaves its locally started gateway and populated dashboard available after the savings report, then shuts it down cleanly on Ctrl-C. Non-interactive runs still clean up automatically, and a demo never retains or stops a gateway it did not start.
Malformed bootstrap administrator identities and weak bootstrap passwords now stop startup before Punk creates an organization or user, while valid local-domain and internationalized addresses remain supported. Home also keeps core run and savings evidence available when optional opportunity or learning-history reads are degraded.
Email verification now completes an already-committed trial when an acknowledgement is lost. Reconciliation is bounded, conflicting audit evidence is refused, and temporary storage uncertainty stays retryable instead of leaving a valid signup stranded.
Broader Compatibility
Typed SDK provider errors now retain a validated dashboard run destination when the gateway supplied one. Malformed request correlation and credential-bearing destinations remain absent, including buffered and streamed provider failures.
Successful typed SDK model results now include the gateway's validated request id across buffered Chat, Responses, Anthropic, and Embeddings calls and every streaming item. Applications can attach the same correlation value to support reports without a separate lookup.
The typed Punk SDK now accepts OpenAI-compatible parallel_tool_calls, seed, frequency_penalty, presence_penalty, reasoning_effort, verbosity, service_tier, logit_bias, logprobs, top_logprobs, prompt_cache_key, and store controls. Buffered and streamed requests preserve the values already supported by the gateway. Legacy user identifiers also pass through, while new integrations are directed to safety_identifier and prompt_cache_key.
OpenAI Responses clients can now send text: null as omission and use bounded top_p, non-negative max_tool_calls, or the documented truncation strategies without forcing live-only routing. Punk preserves native provider values, isolates each explicit value in cache identity, and rejects invalid controls before provider work. Explicit truncation can reuse only an exact provider result; artifacts stay out of that path so they cannot bypass provider context handling.
Research agents can now send tools: null or instructions: null to OpenAI Responses as omitted controls. Punk removes either field before native provider requests, shares exact-cache identity with omission, keeps declared tools and instruction strings distinct, and still rejects invalid populated values before routing.
Research agents can now send metadata: null to Anthropic plus tool_choice: null and reasoning: null to OpenAI Responses as omitted controls. Punk removes those values before native and cross-provider requests, preserves exact-cache identity and token economics, and keeps populated controls on their existing safe paths.
OpenAI Responses clients can now send the official top-level string input shorthand. Punk normalizes it to the standard user message, shares exact-cache identity with the equivalent expanded form, and preserves subject, safety-identifier, and unknown-extension isolation.
Research agents can now omit blank OpenAI or Anthropic user identities plus null Anthropic temperature, top_p, top_k, thinking, service_tier, stop sequences, and stream. Punk removes those values before native and cross-provider requests and treats semantic controls like omission in token estimates and optimization proof, while real identities, sampling values, thinking objects, service tiers, stop controls, and stream flags still pass through unchanged or stay live-only where required. Live embedding responses also fail over before returning incomplete, mis-sized, non-finite, or internally inconsistent vectors.
OpenAI-compatible embeddings now treat null, empty-string, and whitespace-only optional fields as omitted at the raw gateway boundary, and coerce strict digit-string dimensions to bounded integers. The typed SDK applies the same dimension coercion and omits null and empty optionals plus whitespace-only user, so research-agent batches stay aligned without weakening validation of real values.
All eight Dropbox read tools now report post-dispatch transport failures as retryable reads even though Dropbox uses POST endpoints. Share creation, folder creation, and deletion retain explicit outcome-unknown protection and are never marked safe for automatic retry.
The managed Codex launcher and real-client conformance recorder now support image inputs, including inline data URLs from the current desktop client. Promoted text artifacts also leave opaque image and audio interpretation with the live provider instead of answering from placeholder-shaped text.
Artifact and semantic-cache routes now report multilingual output usage from provider-visible UTF-8 bytes across OpenAI and Anthropic responses, so optimized savings match live fallback accounting.
OpenAI Responses clients now recover when their configured Punk base URL already ends in /v1. The recovered path uses the same native authentication errors, validation, rate limits, cache boundaries, streaming events, and run evidence as the canonical endpoint.
Official OpenAI and Anthropic SDKs now receive their native structured authentication errors for invalid credentials and retryable credential-authority outages. OpenAI, Anthropic, and Responses streams also signal compatible reverse proxies not to buffer SSE, reducing time to the first visible event without changing response content.
OpenAI Chat Completions usage now retains documented audio and accepted or rejected prediction token details across live, streaming, and exact-cache responses. Contradictory token aggregates fail before optimized reuse instead of becoming durable billing evidence.
Deterministic OpenAI Responses requests can now reuse exact results within one safety_identifier while different identities stay isolated. Punk preserves opaque safety identifiers byte-for-byte for the provider instead of imposing a gateway-specific length or whitespace contract.
OpenAI streaming callers can now send stream_options.include_obfuscation: false through Punk unchanged. Requests that require provider-added padding still fail before provider work because Punk reserializes SSE and cannot preserve that wire contract.
Required Anthropic service tiers and context management now stay attached to the Anthropic provider during recovery, while documented prompt-cache hints and the default auto tier retain cross-provider fail-open availability in buffered and streamed calls.
Anthropic citation-bearing requests and active OpenAI provider extensions now remain attached to their provider family during recovery, preventing provenance or extension semantics from being silently discarded. Stateless Responses requests, neutral controls, and documented prompt-cache hints still fail over across healthy provider families.
OpenAI Chat Completions requests can now constrain automatic or required selection to a named subset of declared tools. Punk forwards the documented allowed_tools contract byte-for-byte on live requests and rejects any provider response that attempts a tool outside the permitted subset.
Provider-owned continuation state now remains attached to its provider family: non-null OpenAI Responses previous_response_id requests and non-null Anthropic containers fail explicitly when that family is unavailable, while stateless requests retain healthy cross-provider fail-open behavior.
OpenAI assistant audio references now continue through Punk as provider-bound live requests. Opaque provider ids pass through byte-for-byte, contribute only a redacted hash to trace identity, and never enter response caches, artifacts, substitution, or cross-provider failover.
More Trustworthy Runtime Evidence
Valid browser sessions now retain dedicated file-backed authority capacity during unrelated user-read bursts. Home's advisory run totals also use a separate read lane from exact Run Detail evidence, so both login and detail recovery remain available after SQLite contention clears.
Exact-cache tool-turn reuse now requires a provider-valid terminal role and an exact link to an earlier assistant tool call. A caller-owned id field on a user or assistant message cannot make an incomplete transcript look finished.
File-backed organization-member, approval, routable-artifact, cache-inventory, artifact-evaluation, and Governance-total reads now use the bounded database-owned worker pool. Cache inventory stays read-only and lifecycle-bound, while each Governance facet set shares one snapshot. Exact tenant, app, lifecycle, and ordering authority is unchanged; an unavailable optional artifact read falls open to the live provider with explicit route evidence.
First-run, second-run, and sustained-value reports now ignore every request marked with a server-owned request type. New internal gateway work remains visible in the trace ledger but cannot become customer activation proof merely because a blocklist has not been updated.
The seven-day sustained-savings measure now requires verified savings on an input not seen earlier in the same app and agent as first use. Scheduled repeats, context-only changes, and cross-workflow traffic no longer look like durable customer value.
Credential and organization-invitation pages now keep the gateway event loop responsive during slow file-backed SQLite reads. Results retain tenant scope and newest-first order without returning stored credential material or invitation token hashes.
Punk now rejects embedding requests when the selected live configuration cannot serve embeddings instead of returning plausible mock vectors under that provider identity. Explicit offline mock mode and real embedding-capable providers keep their existing behavior.
Owned collection requests such as audit logs, reports, deployments, and tickets now stay on the live provider path across common polite and follow-on wording. Static schemas, fields, documentation, and fixed input remain eligible for safe reuse.
Standalone workers now restore scheduled workflow sweeps for every known tenant after restart. Startup preserves an existing pending, running, or retryable sweep, so recovery does not create a competing schedule execution while an overdue chain is active.
Home, Work, and Results now keep current stored requests, patterns, and artifacts visible when advisory exact totals stall. The API reports those totals as unavailable, coalesces matching reads, isolates tenant scopes, and bounds retained count work instead of leaving the landing pages pending.
Usage & billing now keeps the current plan, spend, and savings visible when quota inventory stalls. Unknown totals are labelled unavailable, while healthy remote-storage reads run together inside the same bounded evidence window.
Common terse requests for live weather, market prices, sports results, service health, balances, and inventory now stay on the live provider. Static explanations, fixed fixtures, and historical weather topics remain eligible for response-cache reuse.
Direct and terse pollen count, index, and level requests now also stay on the live provider. Definitions, historical years, and quoted pollen fixtures remain reusable.
Direct and terse UV index, level, and rating requests now also stay on the live provider. Definitions, historical years, and quoted UV fixtures remain reusable.
Learning now keeps current stored synthesis attempts and candidate conversion rates visible when advisory exact totals stall. The API and dashboard report the total as unavailable, keep recent promotion and live-serving evidence, and bound tenant-scoped count work instead of hiding the attempt log or conversion.
Direct requests for current exchange rates or currency amount conversions now remain on the live provider. Static definitions and explanations of exchange-rate concepts remain reusable.
Common remaining flight-status forms now also stay on the live provider, including flight-status word order, landed or departed checks, and identifier-first shorthand. Static flight explanations and identifier-free questions remain reusable.
Jobs now keep current stored queue rows available when advisory status totals stall. The API reports totals as unavailable, coalesces matching reads, isolates tenant scopes, and caps retained count work instead of leaving the jobs inbox pending.
Audit now keeps current stored events and decision filters available when advisory totals stall. The dashboard reports totals as unavailable, coalesces matching reads, isolates tenant scopes, and caps retained count work instead of leaving Governance pending.
Semantic Web keeps current snapshot rows and Fetch as SOM available when the advisory exact-total read stalls. The Web page reports the total as unavailable, keeps tenant scopes isolated, and bounds retained count work instead of leaving the inventory pending.
Semantic Web also keeps stored snapshots and Fetch as SOM available when open-session inventory stalls. The page reports session evidence as unavailable within a bounded wait, while ordinary slow healthy sessions still appear.
Artifact details now keep proof, promotion-gate evidence, replay, rollback, and quarantine controls available when optional evaluation history stalls. The dashboard shows a neutral unavailable receipt instead of a false empty state, and retained storage work stays bounded by tenant and process.
Billing now renders authoritative usage when optional plan or attribution reads stall. Slightly delayed valid optional evidence remains visible, unavailable sections stay neutral, and authoritative usage retains its full request budget.
Governance now renders required organization, policy, and provider-key controls when optional settings evidence stalls. Run Details likewise preserves authoritative run, trace, side-effect, and feedback evidence when optional ledger-integrity proof is unavailable, without showing an unproven DLP or integrity claim.
The support evidence packet now keeps the run, trace, route, replay, and feedback evidence available when nearby audit history stalls. The packet labels that audit slice unavailable instead of leaving the whole handoff pending.
Home now returns a typed retryable response inside a bounded deadline when its core overview evidence stalls, shares one current read across matching requests, and retries from fresh storage evidence after recovery. Direct questions about current availability, health, balances, prices, scores, and weather also stay on the live path, while explicit static explanations and examples remain reusable.
Repaired live answers and governed web denials no longer wait indefinitely for optional trace or deny-audit writes. Punk keeps the live fallback or provider-free denial authoritative, bounds retained audit work, and omits sensitive URL details from generic evidence. Artifact routes still require both success events to commit before optimized egress; unavailable proof fails open to the live provider.
Punk CLI smoke checks and SDK model clients now accept only current Punk route names and canonical run_ receipts. Proxy-combined, path-shaped, or otherwise invalid headers can no longer become trusted dashboard proof, while valid model output remains available with unknown evidence.
Completed provider answers now stay available when terminal run summaries or failed-artifact route accounting stalls. Rejected cached answers also switch the affected runtime to conservative live routing whenever durable suppression is delayed or capacity-bound, so a retryable feedback error cannot reopen the rejected optimization.
OAuth refresh failures no longer copy token-endpoint response bodies into operator or durable MCP errors. Punk retains the standard recovery code while keeping secrets and hostile upstream text out of evidence.
Concurrent BYOK vault reads now share one bounded lookup, remain capped through cache invalidation, and cannot let a late stale lookup restore an old credential. Tripwire webhook delivery is likewise bounded so a stalled notification cannot delay a provider-free safety block.
Browser-session authority reads now fail closed with protocol-native retry guidance inside a bounded deadline. Concurrent requests share one current lookup, retained generations stay capped, and a recovered authority can serve the same session even if an older read never settles.
Known policy decisions no longer wait indefinitely for their ledger append. Allows continue through the live provider, denies remain blocked, and incomplete policy evidence suppresses cache and learning admission with bounded tenant and process retention.
Aggregate learning now uses one privacy-specific opt-in snapshot per sweep. If that evidence stalls, Punk clears cross-tenant influence, keeps retained reads bounded, and retries from fresh evidence on the next sweep.
Optional DLP and tripwire audit writes can no longer accumulate without limit during a ledger outage. Safe serving decisions remain responsive, noisy tenants cannot consume every retained slot, and capacity returns after storage settles.
API-key serving-mode updates now recover a committed change after acknowledgement loss. If the write outcome cannot be proven promptly, Punk returns a bounded retryable response and keeps the existing serving mode authoritative.
Feedback invalidation and suppression failures now retain only generic durable evidence, preventing backend diagnostics from reaching trace, run-detail, evidence-packet, or replay views.
Tripwire blocks now stay responsive when trust scoring stalls or rejects. Retained work is bounded by tenant and process, late recovery still updates trust, and operator evidence remains generic instead of storing backend diagnostics.
Managed MCP connection pools now stay isolated across server authorities while API and workflow aliases of one server continue to share posture and connections. Durable worker shutdown also reaches cooperative in-flight handlers, preserves retry evidence, and safely tracks an immediate restart; invalid nested Plasmate SOMs fall back to the builtin compiler before evidence can be stored.
Paid gateway answers no longer wait indefinitely for optional trace-integrity telemetry. Punk reports unavailable integrity evidence conservatively, bounds retained reads per tenant and process, cleans up late settlement, and leaves routing and durable ledger proof unchanged.
Overview and the Work dashboard now stay available when historical per-work activity stalls. Operators retain current requests and company-wide results, see a clear warning that historical work totals are incomplete, and recover automatically when storage responds again.
Automatic promotion now holds when Punk cannot read the canary rollout policy and resumes the configured rollout only after recovery. Operator evidence remains useful without copying backend error details into the learning report.
Requests now stop retryably before provider work when the initial run record cannot be confirmed, while late records converge to failed evidence after storage recovers. Preference, open-ended, high-risk, and unknown legacy task classes remain approval-gated even after replay and shadow proof.
Gateway evidence and policy paths now stay bounded when approval or terminal trace storage stalls: callers receive protocol-native retry guidance, providers are not invoked for unresolved approvals, incomplete evidence cannot enter optimization, and per-tenant caps prevent noisy-neighbor failures. Historical tool imports keep the strongest conservative side-effect risk, while clarified current-state requests remain live until a completed answer without sacrificing exact reuse for static workflows.
Serverless ticks now drain already-ready jobs even when a recurring maintenance seed never settles. Punk shares and bounds the stalled work, marks the receipt degraded, and resumes normal seeding after the original operation eventually completes or rejects.
Optional-control trace outages no longer withhold a conservative live answer. Punk bounds unresolved audit writes per tenant and process, preserves the route explanation, and records late audit evidence when storage recovers.
When a provider omits output usage, Punk now estimates multilingual buffered and streamed output from UTF-8 bytes across OpenAI, Anthropic, and Responses. Explicit provider usage, including zero, remains authoritative.
Serverless ticks now continue draining durable jobs when an individual recurring-work seed rejects. The receipt names only the failed seed categories, marks the tick degraded, and retries those categories on the next scheduled call instead of stranding already-ready work.
Current-state requirements in system and developer instructions now bypass response reuse just like equivalent user requests. Explicit static fixtures remain cacheable, so freshness protection does not turn repeatable classification and formatting work back into paid model calls.
When a provider omits usage, Punk now estimates the complete provider-visible request instead of only message text, including tool contracts and native Responses instructions. Multibyte inputs use serialized UTF-8 size, and authoritative provider usage—including zero—still wins.
Exhausted cross-provider failures now leave the run attributed to the terminal attempted provider and model in both buffered and streaming paths. Provider-readiness evidence therefore points operators to the credential that actually failed without inventing a successful servedBy route.
OAuth recovery no longer redispatches an MCP write after an authentication-shaped post-dispatch failure. Punk refreshes the connection for later work while returning explicit outcome-unknown, unsafe-to-retry evidence for the original action.
Workflow runs now snapshot their redaction controls once and apply that decision consistently across trace and governance evidence. Successive-input learning reads trace histories in tenant/app-scoped batches, and pattern inventory facets use one grouped snapshot, preserving measured optimization quality and isolation while reducing storage-bound latency.
Abandoned OpenAI and Anthropic response streams now cancel their established upstream provider work when the downstream reader disconnects. Partial streams finalize as failed and cannot seed cache, evaluation, or learning evidence.
Malformed legacy identity values are now concealed across signed-in account, organization, and audit evidence without rewriting the append-only ledger. Completed live answers also stay available when a failed artifact or learned-route accounting write stalls, and large run inventories compute complete tenant/app-scoped facets in one aggregate read.
API-key revocation now commits together with its audit receipt and reports confirmed, unchanged, or outcome-unknown results explicitly under failures and concurrent requests. Governance screens also conceal legacy identity values that fail current validation—including the signed-in footer—while preserving ID-bound repair and removal actions. Buzz operator evidence loads recent outbox state in one tenant-scoped batch instead of one request per event.
Tripwire inventory and qualified model-substitution cadence reads now share and cap stalled storage work instead of pinning customer traffic indefinitely. Slightly delayed valid evidence still preserves governance blocks and cheaper-model routes; a true outage falls back with explicit route evidence and recovery starts from a fresh read.
App-pinned administrative credentials can no longer enter tenant-wide user and organization identity controls. Tenant-wide administrators and signed-in organization operators retain their existing member, rename, and invitation workflows.
External prompt ingestion and savings evidence now stay responsive and tenant-isolated when storage stalls. Late definitive ingest failures receive one safe idempotent retry, slightly slow valid savings remain available, and Home and Results preserve current request and recovery evidence while naming a savings outage explicitly.
Exact-cache route discovery now reserves retained capacity per tenant during storage stalls, so one noisy tenant cannot turn another tenant's proven repeat into a paid call. Concurrent learning-report reads also share one bounded projection request and retain the latest valid evidence when storage is delayed or rejects.
Route discovery and usage projection now stay bounded and tenant-isolated during storage stalls. A noisy tenant cannot consume another tenant's retained discovery capacity, paid live answers preserve their authoritative status and body, and incomplete usage evidence never enters optimized reuse.
Completed dashboard chat replies now remain available when advisory run summaries stall, and one tenant's outage cannot consume the process-wide projection budget or erase another tenant's real route and cost evidence. Cache-audit retention is likewise partitioned so healthy tenants keep their own cache evidence during a noisy tenant's storage outage.
The newest operator review of an artifact-served run is now the only feedback signal allowed to change artifact confidence and lifecycle evidence, even when an older request resumes through another storage owner after its trace acknowledgement stalls.
Retried or overlapping model-substitution shadows now contribute one immutable quality sample and one route statistic. Anthropic prompt-cache writes and reads are reflected at their TTL-specific list prices in stored cost and savings evidence, including compatibility with nullable provider counters.
Paid live answers now finish within a bounded evidence window when best-execution or post-live shadow storage stalls. Punk marks incomplete or indeterminate evidence explicitly and prevents it from entering cache, shadow, or learning reuse until proof is complete.
Buffered OpenAI Chat Completions can now return requested token log probabilities through Punk's live path; mismatched token-byte evidence is rejected and DLP changes suppress stale confidence metadata. Paid buffered answers also remain available when completion evidence storage stalls, with tenant-scoped retained-work limits, while cache hits require the complete normalized wire snapshot to remain unchanged before serving.
Provider readiness now distinguishes configured credentials from recent live gateway proof, ignores imported or cross-tenant evidence, and gives operators a concrete recovery action after stale or failed traffic. Provider-key replacements preserve the working credential on failure and commit rotations with their exact audit evidence; password changes invalidate every other browser session.
Delayed artifact inventory and preference-evidence reads fail open within a fixed bound, cap retained work across repeated bursts, and revalidate lifecycle state so a stale promoted snapshot cannot route after quarantine.
Authenticated traffic and governance denials now remain bounded through credential-index, memory-influence, and tripwire-evidence outages. Unavailable key authority fails retryably closed without provider calls, memory quarantine retains approval-required behavior with capped fallback work, and detected tripwires stay hard-blocked while advisory event writes settle.
Home remains usable when background-job telemetry is unavailable, usage reports attribute complete requested windows without double-counting concurrent activity, and cache hits or live fallback stay responsive during delayed audit writes while recovered evidence still settles durably.
Large run evidence packets now download completely even when they exceed the dashboard's interactive response limit. Cache routes report savings and avoided latency from the observed source call—including legitimate zero-cost and zero-latency evidence—instead of substituting estimates. Provider-affine OpenAI requests now fail explicitly when their required provider family is unavailable, rather than returning a simulated mock answer; ordinary non-affine requests still retain healthy live-provider failover.
Complete SDK evidence exports now stream with backpressure, cancellation, and independent response limits, so slow consumers can retrieve large packets without unbounded buffering. Optimization proof now includes sampling controls; existing proof safely falls back live once and re-earns evidence instead of serving an answer produced under different sampling behavior.
Gateway failures now include browser-readable request ids that typed SDK errors preserve for support correlation. Anthropic artifact routes conservatively honor output caps and fail open to the live provider, scheduled worker ticks reject unknown or ambiguous controls before starting work, and OpenAI SDK streams request usage while retaining complete terminal model, provider, and token evidence across live and optimized routes.
Semantic-answer reuse now partitions deterministic values across every conversation role, Anthropic SDK streams preserve monotonic token and cache-accounting evidence, positive SDK notes cannot become corrective failures, and usage reports reject ambiguous windows instead of silently changing them.
July 2026
A Faster Path To Measured Value
Repeated structured JSON work can now compile text fields assembled from several request values into deterministic templates. Punk admits only one bounded skeleton proven across training and holdout traces; inconsistent or runtime-reserved values stay on the model-backed path.
- The Overview now leads with what needs attention, the next recommended action, and progress from first traffic to a proven optimization, then asks operators for the safe repeat that demonstrates reuse. First-time operators also get an app-specific connection receipt that distinguishes received traffic from a served run without exposing sibling applications; SDK applications can submit typed named metrics against settled runs.
- Trial, onboarding, agent activation, and design-partner paths are clearer: the offline demo starts and cleans up an isolated mock gateway and waits for its complete shadow batch, generated starters preflight file collisions and verify fixed-task answers before recording positive feedback, hosted smoke tests separate required route proof from optional savings, expose observe-mode would-have savings without claiming they were served, reject inconsistent exact replay, and keep onboarding evidence shareable. Codex diagnostics now distinguish a broken installation from an unsupported version and provide a direct recovery action, while the offline demo names the managed authority required for production tool-result reuse.
Proof That Is Easier To Inspect
- Run, pattern, and artifact views explain routing decisions in plain language while keeping replay, shadow, policy, and fallback evidence available for deeper review; app-scoped run evidence is filtered before bounded result windows, concurrent replay requests converge on one immutable proof, counterfactual audits cannot mix another tenant's evaluations, and ledger proof distinguishes complete hash chains from compatible partial legacy coverage. SDK clients now expose and validate complete route explanations and memory-influence acknowledgements against the requested run, retain server-owned tool-cache ineligibility evidence, and can correlate exact LangChain response headers through a failure-contained
onRunobserver. - Punk published a measured public reference study covering 150 benchmark requests and 50 held-out treatment requests. It is explicitly presented as mechanism proof—not customer or production evidence.
Safer Production Behavior
File-backed SQLite Runs inventory reads now run outside the gateway event loop. Slow inventory storage no longer delays unrelated health or serving requests, and a shared worker limit keeps concurrent row reads within a fixed runtime budget.
Concurrent administrators can no longer remove every owner from an organization, even when requests run through separate storage instances. Production readiness checks now settle with explicit dependency timeout evidence and share stalled probes across polling bursts instead of retaining one storage read per poll.
Concurrent workflow and agent creation now reserves plan capacity atomically across runtime instances, while validated API-key and browser-session traffic stays available through stalled activity bookkeeping without retaining one write per request.
- Cache invalidation now commits with its operator receipt or rolls back, including recovery from a lost storage acknowledgement without purging newly created entries. Legacy clients can request fresh work with
Pragma: no-cachewhen modern cache controls are absent, and MCP registry mutations reject misspelled or server-owned fields before applying configuration. Invalid Stripe, Chargebee, and Shopify financial-action caps fail closed before level-4 plans; concurrent dashboard turns retain complete alternating history; concurrent replay requests commit one proof; abandoned inbound bodies settle when callers disconnect; and recovered Anthropic/v1/v1paths preserve canonical validation-first budgets. Shopify, Twilio, and Mailchimp reads now stop stalled or oversized upstream responses while preserving successful JSON that arrives without a media-type header; accepted oversized connector writes retain request identity and are explicitly marked outcome-unknown and unsafe to retry. Optimized routes fall back to the configured live provider more consistently when evidence is missing, stale, inconsistent, or unsafe; healthy provider backups remain reachable through rejected or stalled failover-audit writes; hard MeshGuard denials override approval hints, observe-mode SDK writes execute with audited evidence while optimize-mode gates remain enforced, provider refusal/content-filter terminals remain successful but non-reusable across buffered and streamed protocols, throttled chat requests retain blocked-run explanations, undeliverable team invitations are revoked with their denial audit, live answers remain available during request, terminal, model-completion, or run-projection evidence outages without seeding future optimization, provider-requested storage stays live and bound to the selected provider, validated keys and sessions survive activity-timestamp outages, canary promotions pause when rollout policy state cannot be read, and opportunity estimates use retained 30-day live work rather than stale lifetime bursts. - App and tenant boundaries, approval flows, audit reasons, permission-fingerprinted tool caches, and high-impact lifecycle controls are more conservative and easier to verify; app-pinned administrators now see only their app's agent identities and live web sessions, see and decide only their app's approvals, and cannot inspect or mutate tenant-wide settings, credentials, connector setup, MCP-server configuration, or tenant-wide tripwire controls, individual API keys can move atomically between optimize and observe modes, and administrators can recover passwords through enumeration-neutral reset links that invalidate prior sessions.
- Runtime telemetry, SLO checks, load testing, bounded provider, MeshGuard, connected MCP, Slack/shared connector, SDK, SEC research, and GitHub waits and response bodies, private SDK no-store requests, post-header cancellation, usage-accounting fail-open continuity, atomic recent-window dashboard chat turns, outcome-unknown non-retryable write timeouts, write-aware scheduled retries, redirect-contained workflow notifications, retryable MCP OAuth recovery, and dead-lettered crashed leases make production readiness and operational risk clearer; non-cooperative workflow writes cannot overrun their deadline, bounded scheduler stalls no longer duplicate live jobs, MCP explanations retain readable rejected-route evidence, Stripe setup contracts fail before success, retention removes run evidence atomically, confirmed email delivery survives stalled response cleanup, and the dedicated Buzz worker verifies an already-migrated schema without DDL and starts commissioning before its live evidence becomes the release-readiness gate.
Abandoned dashboard chats and manual workflow or agent runs now stop downstream provider work instead of continuing after the caller leaves; onboarding distinguishes served traffic from side-loaded history, and large run timelines stay within the dashboard response budget while retaining complete export evidence. Tenant settings and opened web sessions now commit with exact audit proof or roll back; cancelled MCP writes retain secret-safe outcome-unknown/no-retry evidence; and durable webhook retries expose stable bounded keys for receiver deduplication. Transient run-start writes now recover without duplicate paid work, while rejected or stalled model-request and DLP audit writes stay bounded without blocking healthy providers or truncating already-masked answers. Negative reviews now remove the affected artifact from routing immediately, including after a retried lifecycle fault, ambiguous workflow names require an explicit workflow id instead of silently choosing a target, incomplete provider outputs stay out of learning and replay proof, artifact evidence honors trace redaction without changing served answers, and Chorus output tripwires inspect assistant text plus tool arguments while retaining honest incurred-cost evidence. SDK side-effect traces retain idempotency evidence through terminal projections without hiding distinct tool executions, and Linear calls stop stalled or oversized responses while keeping retries read-safe and timed-out writes outcome-unknown.
Broader Compatibility
Anthropic citation-enabled document requests now retain document, web-search, and custom search-result citations through buffered and streamed live responses. Citation metadata is bounded, covered by DLP and tripwire enforcement, accepts nullable provider titles and stable non-URL retrieval identifiers, and stays outside response caches.
OpenAI streaming responses now preserve requested token log probabilities with response-wide bounds and correct reconstruction when one Unicode character spans adjacent token-byte arrays. Copied base URLs that already end in /v1/chat/completions also recover without producing a duplicated /v1 route.
Buffered OpenAI responses now preserve validated URL citation annotations and exact Unicode spans on the live path. Citation metadata is covered by DLP and tripwire enforcement and stays outside response caches so provider provenance cannot be replayed as local evidence.
Current OpenAI and Anthropic browser SDKs can call Punk directly with their official metadata headers, and SDK run inspection now exposes the latest byte-bounded event window instead of failing once a valid trace exceeds the client's response ceiling.
- The integration catalog now includes 51 managed SaaS connectors and 517 governed tools across engineering, support, data, finance, security, and business systems.
- MCP setup now accepts copied OpenAI and Anthropic endpoint URLs while preserving nested reverse-proxy roots, and the Runs view can filter exact serving-provider evidence across scoped pagination. OpenAI and Anthropic compatibility now includes native Claude image blocks with canonical fallbacks for unsupported references, native tool-call streaming, safe cross-protocol streamed-plan reuse, stop-control preservation in both cross-provider directions, matched Anthropic stop-sequence evidence across live, cached, and streamed responses, serial tool-call intent across buffered and streamed failover, cache isolation for distinct OpenAI token-limit controls and ordered tool catalogs, live-sibling failover before mock fallback, cancellable and output-bounded provider streams with prompt upstream release, runtime rejection of invalid declared text without breaking empty tool-only responses, independent bounds for oversized provider bodies and non-cooperative polling transports, surfaced completion/stop reasons, complete bounded streamed tool-call assembly, buffered and streamed refusal fidelity with explicit SDK safety-stop evidence, signed Claude thinking continuity across reusable tool plans without learner exposure, served-model fidelity across live/failover/cache paths, deterministic repeat-cache defaults in generated Anthropic starters, preserved assistant refusal history and provider backoff hints, bounded stalled-stream recovery, streamed OpenAI usage/completion proof, streamed reuse of eligible buffered exact responses, and standard
Cache-Control: no-cache,no-store, and zero-age revalidation controls for fresh live work. - OpenAI, Anthropic, OpenRouter, Vercel AI SDK, LangChain, and Claude Code adoption paths now have clearer setup guidance, safer launch validation, examples, and compact SDK route proof with fallback for older gateways; TypeScript SDK tool continuations and multimodal text/image/audio/file inputs compile without casts, cross-provider
top_pstays intact, and native nullable container, context-management, and service-tier response state survive live, exact-cache, and synthesized Anthropic streams while stateful containers remain live-only.
Structured OpenAI output contracts are checked before buffered or streamed answers are served, generated tool schemas can use bounded local JSON Pointer references (including URI-encoded fragments), validated json_object requests can use promoted artifacts without a provider call while non-object results fail open, and Google Workspace calls now use bounded deadlines while large document exports retain a safe truncated prefix.
Standard OpenAI completion metadata now passes through byte-faithfully, follows provider character limits, and keeps those requests on live routes so cache reuse cannot erase provider-visible metadata. Repeated OpenAI streaming requests can now receive eligible semantic-cache answers as valid SSE without another model call, current OpenAI safety identifiers isolate optimized reuse, and requests using different Anthropic API versions no longer share cached response bytes.
Easier Daily Operation
Workflow editors now keep Save and Run available when optional savings, recent-run, or MCP inventory evidence stalls. The editor renders within one bounded wait and labels unavailable evidence without hiding the workflow.
- A keyboard command palette, better linked evidence, clearer filters, prefix-scoped cache eviction, safer lifecycle actions, SDK rollback/quarantine and trustworthy thumbs feedback, and resumable ignored patterns reduce the effort required to navigate and operate the runtime.
June 2026
Adaptive Runtime Foundation
- Punk launched with OpenAI-compatible and Anthropic-compatible gateway surfaces, policy enforcement, route explanations, provider failover, trace evidence, and cost tracking.
- Repeated work can move through evidence-gated cache, model substitution, and deterministic artifact routes, backed by replay, shadow, canary, rollback, quarantine, and approval controls.
Operating Surface
- The initial dashboard brought together runs, patterns, artifacts, learning, savings, governance, workflows, agents, and approvals, with hosted organizations, API keys, usage metering, and billing support.