Community group telemetry backend
Scope
This backend implements the aggregate-only slice of epic #176. A VRDex-owned VRChat account observes a connected group's member count and visible group instances. It never accepts a customer's VRChat credentials and does not store instance user lists, usernames, or user IDs.
Control plane
The club workspace separates action permissions from analytics visibility. Connection actions require manage_integrations; event-association actions require manage_events. Private dashboard reads admit owners and staff, then filter each data category. Publication settings are owner-controlled.
communityVrchatIntegrations is the community-owned lifecycle record. Connect allocates one healthy account only while assignedGroupCount < capacity - reservedHeadroom; Convex mutation serialization prevents concurrent over-allocation. The same VRChat group cannot be active on two community profiles. Disconnect immediately disables collection and public fields, then a fenced disconnecting assignment makes the service account leave the group before releasing account capacity. Reconnect resets freshness and opens a new telemetryEpochStartedAt; private and public projections filter observations, sessions, coverage, associations, and rollups to that epoch so a previous group's retained history cannot appear under the new connection. Ingestion also refuses to reuse an open session from an older epoch.
collectorAccounts stores an opaque alias, VRChat service-account ID, capacity, health, request budget, credential generation, and an external secret reference. It never stores the credential. Local proof authentication reuses an alias-scoped operating-system vault session only after validating the authenticated immutable account ID; invalid sessions are removed and there is no plaintext fallback. The production worker resolves its provider session from the external secret. That gate was cleared by BASIC's 2026-07-27 risk acceptance of durable VRChat service-account sessions for VRDex-owned proof accounts — a product-owner decision, not a VRChat grant, and the stop condition in docs/planning/community-group-telemetry.md still stands. The secret must record the vrchatUserId it belongs to, and the worker sends it on every control-plane call so a collector paired with another account's secret cannot do work. collectorAccountLeases grants a bounded work claim. communityVrchatIntegrations.leaseGeneration makes fencing tokens monotonic across release and reassignment, so an old worker cannot resume writes with a reused token.
The worker claims one integration immediately before collecting it, so a slow poll cannot age leases for other queued integrations. A one-claim request skips due integrations leased by another worker and continues to the next eligible group. Among due groups it serves the least recently claimed first, preventing a continuously due management queue from starving its neighbors. Shutdown or heartbeat failure before collection releases the new lease. Account and integration request budgets still govern each provider call.
Reassignment is an internal operator action and is allowed only after the source account is quarantined, retiring, or retired. The mutation serially checks target headroom, releases the old lease, moves the allocation, and opens an unknown coverage window before the next fenced claim. Registering new credentials does not automatically reactivate a quarantined or retired account; an operator must reconcile its external group memberships before explicitly returning it to ready.
The account-specific /telemetry/worker HTTP action authenticates a SHA-256 worker key using a constant-time comparison. Every non-claim operation also binds the active lease to that authenticated collector account, worker ID, and fencing token. Lease expiry is checked against trusted server time; collector observation timestamps are validated separately within a bounded clock-skew window. Runtime liveness uses heartbeat, carrying the exact release SHA, bounded version, capabilities, and consecutive control-plane failure count. Telemetry uses claim, membership result, aggregate ingest, failure, and release. Ownership proofs use proof_claim, proof_budget, proof_result, proof_outcome, proof_release, proof_auth_failure, and proof_rate_limit. Proof operations are not lease-scoped. They are bounded by the account's request budget instead, of which proofs may take min(floor(limit / 2), limit - 2) per minute so an atomic telemetry poll always fits, enforced by shared proof:account:<id> and proof:global counters rather than per worker. The worker key is re-checked after the request body is read, and again inside the mutations that grant ownership or change account state, so a rotation cannot be outrun by a slow request. Responses and operator-health queries omit customer, provider, credential, proof-code, and target identifiers.
Proof dispatch and provider checks are separate lifecycle facts.
lastCheckedAt remains the legacy queue cooldown stamp;
firstDispatchedAt/lastDispatchedAt and dispatchCount describe work handed
out, while firstCheckAt/lastCheckAt, checkCount, and the bounded
lastCheckOutcome describe provider requests that actually ran. An idle worker
stamps lastProofPollAt and its lastProofPollReleaseSha only after the account
and fleet proof gates pass. This lets the claim action distinguish a live proof
reader from a generic running task, while the deployment gate can also reject
an obsolete release.
Observation model
communityPopulationObservations: one immutable total population, instance count, and world distribution per successful poll; the poll ID is the idempotency key.instanceSessions: first seen, last seen, and confirmed close for an immutable provider location. The canonical location combines world and instance identifiers, so the same instance suffix can exist in different worlds. Only successful complete instance enumerations increment the miss counter. Two consecutive misses close a session; seeing the same provider location later opens a new session.lastObservedAtrecords actual visibility, whileclosedAtrecords the later confirmation poll.instancePopulationObservations: aggregate population per visible instance per successful poll. These support exact confirmed-event recaps without collecting people.communityMemberCountObservations: a row on count change or after a six-hour heartbeat.communityTelemetryEventRecapJobs: a temporary cursor and metric accumulator for a confirmed event. Each internal mutation consumes at most 100 instance observations, resolves their current session associations, and schedules the next page. The existing recap remains in place until the complete replacement is ready. Event-boundary changes and new confirmations cancel stale jobs.collectionCoverageWindows: observed, estimated, stale, unknown, or degraded intervals. Missing time never produces a zero observation.
Every observation carries source, collector version, observed time, coverage, and fencing token. The v1 source is first_party; vrcpop and vrcx are reserved adapter values, not active integrations. The adapter replaces subjects embedded in hidden(...) or private(...) instance-locator markers, including legacy user IDs without a usr_ prefix, before ingestion. Defense-in-depth validation rejects unredacted subject markers, remaining usr_ identifiers, foreign group markers, inconsistent world/location pairs, negative/non-integer counts, duplicate provider locations, malformed world IDs, oversized values, and control characters.
Private instance freshness and history
Stored session state records the observed lifecycle. An open row does not establish current liveness. Private list and direct-detail queries separately project server now and nullable liveObservedAt. Live requires an active or degraded integration with its kill switch off, analytics enabled under the existing legacy default, and a current-epoch open session actually observed no later than server time and no more than six minutes ago. Exactly six minutes remains live; one millisecond later does not. A newer non-observed coverage transition revokes the claim even if a later group-only success did not see that session.
Repeated failures with the same state and reason coalesce into one coverage window and advance its updatedAt. Session freshness uses that latest evidence, not only the window's start. Closing a non-observed window overwrites updatedAt with its recovery time, so the projection conservatively requires a session observation at or after that recovery boundary. A late observation inside the gap remains in history but cannot restore Live, even if it postdates the last failure. A session actually observed at recovery or afterward can restore Live within the normal freshness lifetime. A group-only recovery never refreshes the session itself.
The telemetry Live list filters that projection. Instance history and Home Recent instances include all authorized current-epoch sessions in descending openedAt order through the existing community/time index. Fresh rows deliberately appear in both lists. Filtering preserves bounded pagination and continuation through empty pages. Session state, close timestamps, raw population and statistics remain intact after outage, disconnect or analytics disablement. An unconfirmed close displays Unknown once freshness expires; a confirmed close keeps its recorded timestamp.
Each list or detail query owner supplies a fresh mount/scope nonce. The shared display timer fixes the initial server/monotonic calibration and expires against each session's original observation deadline, without database writes. Reactive evaluations cannot renew the same observation. Detail expiry unmounts an open closure confirmation. Provider-live management keeps its independent permissions and separate provider-read lifetime.
All rows in a paginated list share that one owner calibration, including later pages and newly observed rows. Each list owner calls the category-authorized getInstanceListClock query with the same fresh nonce as its pages. It returns only server time, including when the first filtered page is empty, so later rows cannot restart pagination. Each observation is validated against its own row's server evaluation time, then expires against the original calibration. A later observation may postdate the original clock; repeated evidence never renews its deadline. Clock response time counts against the available lifetime. The shared plural helper schedules only the next unexpired deadline; the existing singular consumers delegate to the same implementation.
Private membership chart coverage
Membership charts use compact successful-poll coverage, independently of the sparse count-change/six-hour heartbeat rows. New observed coverage windows record observedThroughAt. Explicit non-observed transitions end continuity immediately. A successful-poll gap greater than ten minutes closes the earlier observed window at its last successful poll plus ten minutes and starts a new window. The bound allows the five-minute quiet poll delay, a four-minute management pass before aggregate collection, and one minute of dispatch/transport delay. It is a conservative continuity inference, not a collection SLA. Exactly ten minutes remains connected; longer silent outages are blank. Late timestamps cannot rewrite newer collection-state evidence. Population rollup interpolation retains its separate five-minute rule.
The group-size-authorized day query returns only observed interval boundaries and completeness, with no population values, source identifiers or operational reasons. It reads at most 1001 compact coverage rows plus one same-epoch boundary row for a local day of up to 26 hours. Above 1000 rows it returns incomplete coverage and no inferred intervals. It does not read large population/world-distribution records. The day chart keeps actual member observations and inserts null boundaries across gaps. The range chart isolates partially observed days, leaves unobserved days blank, and only carries a prior count into a day with continuous poll coverage. Isolated real observations remain visible as points. Each bucket includes its server evaluation time and an absolute continuity deadline for an incomplete current day. The range query owner uses a fresh nonce per mount/range and the existing monotonic display timer, so a silent outage disconnects the line at the original deadline without another collector write. Reactive reevaluation does not extend that deadline. Fully covered historical days have no expiry; the recorded member value remains visible when current continuity expires. Both consumers retain their existing membership line and independent interactions.
Legacy windows lack evidence of internal silent gaps. They remain retained, but cannot establish continuous membership coverage; their exact member observations stay visible without inferred connecting lines. The first new successful poll begins a new proven window. No historical backfill runs in this change. This conservative legacy display limitation remains until a separately approved bounded historical reconstruction exists.
Rollups and retention
Primary active VRChat group links can collect aggregate member counts without a connected telemetry integration or service-account membership. The worker claims only publicly surfaced community links whose Group size audience is Public. It reads one group through the authenticated account's existing low-priority metadata budget after proof and connected telemetry work. A linked group's first eligible read is due immediately; later reads are due daily. Multiple profiles linking the same group share one group metadata record and count series. Connected aggregate polls write the same series from their existing group response, without another provider request.
vrchatGroupMemberSnapshots stores vrchatGroupId, memberCount, observedAt, and optional groupCreatedAt, indexed by group and observation time. A new row records a changed count or a daily heartbeat. vrchatGroupMemberMetadata retains the validated creation time, latest observation time, and a short claim lease. Both tables retain observations permanently. Provider failures preserve the latest count and timestamp, schedule a shared later retry, and never insert zero. Link removal or a different primary group stops collection through that link without deleting shared history. The public projection applies each profile's own Group size visibility separately.
community-telemetry-v1 rollups use UTC hour/day/event windows and trapezoidal integration between observations no more than five minutes apart. They include current population, active instances, peak concurrency, player-minutes, coverage ratio, member count/growth, and world distribution. Re-running the same window updates the existing versioned row, so late or corrected raw observations deterministically replace the rollup.
The hourly Convex cron schedules the previous hour, previous UTC day, and paged recent-event work. Recent-event work starts private time/world suggestion scans and recomputes every in-window event that has a confirmed session association. Suggestion scans are paged and bounded to the event window plus a six-hour setup lead. Manual confirmation and suggestion approval schedule the event rollup immediately; removing the final confirmed session deletes the event recap instead of retaining an empty public artifact. Raw group and instance observations, session boundaries, coverage windows, member changes, and rollups are retained permanently. Rollup computation does not delete its source observations. There is no age-based telemetry compaction job; deletion workflows are deferred.
Deployment removes the previous daily compaction registration and its scheduled-continuation functions. Until that backend release is deployed, existing deployments retain their previous policy. Data already removed by compaction cannot be recovered by this change.
Public projection
Category visibility takes precedence over legacy publicMetrics booleans once a visibility document exists. Without one, the existing booleans remain the fallback. Group size and membership movement independently gate member counts and net growth, including fields nested inside population history and event recaps. Net growth is not a count of joins or departures.
getPublicCommunityTelemetry is the public instance-analytics projection. An explicit enabledFeatures list without analytics returns no public telemetry, including retained hourly history, instance sessions, member counts and growth, and event recaps. An absent list retains legacy analytics-enabled behavior. Authorized owners and staff can still read retained private history. Each public metric defaults off and is included independently when analytics is enabled. Hourly history is one deliberate bundle containing its documented rollup fields. Public instance history returns at most 20 newest current-epoch sessions when instance_history is public. It includes first/last observed times, confirmed close time when present, and a linked world name only for published worlds. An open session is historical evidence, not a live claim. It excludes provider locations, instance/group IDs, raw observations, population, and unpublished world details. Current population disappears after six minutes without a successful poll, which covers the five-minute healthy quiet cadence plus scheduling tolerance; the projection remains stale if another enabled historical surface is still present. Disconnect returns no public telemetry.
getPublicGroupMembership projects the primary linked group's timestamped latest count, creation time when known, and at most 500 observed points. It works without a telemetry integration and requires the profile's Group size audience to be Public. It combines shared group snapshots with current-epoch connected observations, without exposing group IDs or member identities. Long histories use at most eight indexed sampling windows per saturated source and 32 indexed existence probes for ambiguous multi-day spans. The projection marks a span sampled only when an omitted observation is known, so the chart does not label retained measurements as unobserved or a real missing day as sampled. An integration without an active primary link can still project retained connected observations, but standalone polling requires a primary link. Restoring a public community resets its primary link's poll deadline so a previously deferred first read is due immediately. The public page has separate Appearance switches for the count and history graph; these switches change presentation, not API authorization. A dotted chart segment marks time between creation and the first observation or between raw observations separated by an unobserved UTC day.
profiles.getPublicBySlug attaches this same projection to the community profile. The normal web route, /api/v0/communities/{slug}, /api/v0/profiles/{slug}, hosted MCP, and stdio MCP therefore share one visibility boundary and one PublicCommunityTelemetrySchema. If the connected analytics group differs from the active primary profile group, public telemetry omits the connected group's member count and growth, including rollup membership fields, while preserving its other public instance metrics. This prevents secondary-group counts from being presented as the primary group's membership. Existence-only callers opt out of the telemetry fanout. Internal observation IDs, integration/account IDs, raw coverage reasons, group IDs, and service-account metadata are excluded.
Verification
Backend tests cover authorization, public-off defaults, concurrent capacity allocation, quarantine/reassignment, fleet stops, monotonic fencing, idempotency, concurrent sessions, close/reopen behavior, malformed input, 401 account isolation, redaction, stale public behavior, deterministic recomputation, and preservation of raw observations older than 90 days during hourly and daily rollup recomputation. Worker tests cover the provider adapter, aggregate-only projection, account-scoped operating-system vault records, immutable identity validation, expired-session removal, transient validation failures, slow-metadata caching, request budgets, cadence jitter, 429 backoff, gap-aware player-hours, and diagnostic redaction.