团队文章权力与治理

When registry data changes, why can a working network disappear?

Follow a changed record through route filters and RPKI to see why access can fail, and why Lu Heng argues for continuity and a real path to independence.

目录

Imagine your service still works on a mobile connection, but customers on another network can no longer reach it. The servers are healthy. The cables are connected. One possible explanation is that another network has stopped accepting the route to your addresses after a record changed.

The missing link is the decision between the record and the router. A database change can affect connectivity when a system uses that change to build a filter or evaluate an announcement. To understand the failure, follow that chain. “The registry is wrong” is only the beginning of the diagnosis.

Two buildings remain wired to a server; one network device has dark indicators while a technician examines a card file.
Physical connections can remain intact when a record change affects route acceptance. Follow the evidence and the receiving policy.

First, ask which record changed

An IP prefix is a block of addresses. An autonomous system number, or ASN, identifies a network exchanging routes with other networks. BGP is the protocol those networks use to announce which destinations they can reach.

Several kinds of records sit alongside that exchange. They answer different questions:

  • Registration records: which organization is recorded against an address block, and whom to contact. The RIPE Database documentation distinguishes resource registration, routing information and contact records even though they share a database.
  • Internet Routing Registry (IRR) records: information operators can use to construct lists of routes they will accept. A record describes intended routing; it does not itself announce a route.
  • Resource Public Key Infrastructure (RPKI): signed route-origin authorizations, called ROAs, from which validators produce data for checking the originating network and permitted prefix length.

An outdated contact email does not directly withdraw a BGP route. A missing route in a provider's generated filter can stop that provider accepting it. Calling both “invalid registry data” hides the difference that matters.

How a record becomes a rejected route

Consider a provider that builds its customer filters from IRR data. Here is a possible failure sequence:

  1. A route record is removed, or the customer's routing set no longer includes it.
  2. The provider's next filter build omits that route.
  3. The new filter reaches the provider's routers, which reject the customer's announcement.
  4. If no usable alternative remains, people depending on that path lose access.

The data must actually feed that filter for this chain to occur. NTT DATA's published routing-registry policy provides a concrete example of IRR-based customer filters and automated updates. It also documents rejecting RPKI Invalid routes and suppressing conflicting IRR records. These are identifiable operational mechanisms, not a universal registry off-switch.

What RPKI “Invalid” actually means

For origin validation, the route being announced is compared with validated authorization data. RFC 6811 defines three outcomes:

  • Valid: at least one covering authorization matches the origin ASN and allows the announced prefix length.
  • Invalid: covering authorization data exists, but none matches both conditions.
  • NotFound: no authorization data covers the route.

For example, a route announced by network A can become Invalid if its matching authorization disappears while a covering authorization for network B remains. If no covering authorization remains, the result is NotFound instead. “Invalid” here is a routing-validation result, not a verdict on ownership or institutional legitimacy. Origin validation also does not authenticate the whole path a route claims to have taken.

The next step is the operator's configured policy. RFC 8481 separates setting a validation state from acting on it: policy must be configured by the operator. A network configured to reject Invalid routes can then reject this announcement. A registry edit and a rejection are connected through that machinery; they are not the same event.

Why some people lose access before others

Networks do not all use the same filters, providers or data at the same moment. RFC 7115 explains that RPKI caches can hold different views and that updates have no single guaranteed arrival interval at routers.

In the opening example, one network may already have acted on changed data while another still has an accepted path. That makes a partial outage possible. It does not prove the registry caused any particular outage: engineers still need to inspect the affected route, the data in use and the decision that rejected it.

Find the broken link before changing more things

The useful recovery question is specific: which system stopped accepting which announcement, and why? An operator can work through it with the provider:

  1. Identify the announcement. Record the affected prefix and origin ASN, and check whether the expected route is still being advertised.
  2. Locate the rejection. Compare the provider's installed filter and validation result with the route. Ask which data source and update produced that decision.
  3. Preserve the evidence. Keep the relevant records, observations and timestamps so that a mistaken update can be distinguished from an unauthorized announcement.
  4. Repair and verify. Coordinate the precise record or configuration correction, confirm it has reached the consuming systems, and test access from the networks that failed.

Disabling validation everywhere would also remove protection against bad announcements. The operational task is to restore correct evidence and intended routing, then verify the result. A successful database edit alone does not demonstrate that customers can connect again.

The deeper issue: useful evidence can become concentrated power

This is where the technical chain meets Lu Heng's argument. An administrator need not forward anyone's packets to exercise influence over their reachability. If other systems depend on records it controls, its decisions can impose costs on operators and customers far beyond its office.

In Note 65, Running-Code Primacy, Lu Heng argues that coordination must be limited by the needs of running networks. A role in maintaining records cannot expand itself into a permanent political mandate. The fact that a record is signed answers a technical question about the assertion; it does not establish a right to govern everyone affected by it.

His proposal goes beyond a better complaints process. The common rules should protect uniqueness, verifiable control and interoperability, while participants validate state locally and decide which compatible changes to adopt. A new committee with an unlimited veto would reproduce the problem under another name.

Continuity needs a way out that other networks can use

A usable alternative must carry forward trustworthy records, security assertions and working relationships. Other networks need to be able to verify those records and use them. Simply copying a database does not make them accept a replacement, and declaring independence does not repair a rejected route.

Note 72 makes the intended outcome explicit: accurate records, operational continuity and a real ability to leave a failing coordinator. The practical challenge is to build that ability without losing the shared uniqueness and compatibility that let independent networks communicate.

The guide to registry-state export explores the record-carrying part of that work. It is one component of a transition, alongside verification, operator adoption and continuity testing.

Prepare while the service is still working

Ask your team to trace one important address block from its records to the providers and customers that depend on it. Who can change each record? Who consumes it? How would an error be corrected if the current administrator were unavailable or disputed?

Those answers take time to establish across organizations. Finding them before a failure creates room to act; discovering them during an outage leaves customers carrying the cost. The urgency is to build practical independence while continuity can still be protected.

Continue with Note 65: Running-Code Primacy to follow Lu Heng's full argument for coordination grounded in working networks, local verification and voluntary adoption.