Catching Schema Drift at Service Startup
This is a guest post from Matthaios Stavrou, Staff Software Engineer – Manager at PwC
Apache Kafka® messages are just bytes. Imagine one service publishing an order with an amount as a number while another expects it as a string. Both services might be running, but they disagree on how to read the event. A schema defines the fields and their types so teams can review and test changes to that contract.
Confluent Schema Registry stores those schemas under subjects, versions them, and checks configured compatibility when a schema is updated. It can reject a change that would violate the subject's compatibility policy.
This post looks at two points where application teams can check their own assumptions. CI tests a proposed local schema against the registered contract while the change is under review. Startup checks the registry configured for the deployed service before it becomes ready.
Why check again at startup if CI passed? Between the build and deployment, another team can register a new schema version, change a subject's compatibility setting, or point the service at a different registry. The CI result describes the registry it checked at build time; it cannot guarantee the state the service will see later. A startup check can catch that difference before the service becomes ready.
Accompanying this blog post is an example implementation and demo built on Spring Boot to show the startup checkpoint in a real service; the same delivery pattern can be applied outside Spring.
Start with the contract, not the framework
Confluent Schema Registry already provides the core mechanics for controlled schema evolution. Schemas live under subjects, compatibility can be configured globally or per subject,1 and client applications use schema IDs during serialization and deserialization. A schema ID identifies a registered schema; a subject version numbers its registration under that subject.
A compatibility check tells us whether a schema is allowed under the configured compatibility level. For a live production system, the more practical problem is when the engineering team finds out that it is not.
For most engineering teams, the earliest useful place is the pull request or CI build.
Schema Registry still owns the compatibility decision. The workflow just brings that decision into the development path sooner.
Make incompatibility a build failure
For JVM teams, compatibility checks fit naturally into Maven- or Gradle-based builds. The Schema Registry Maven plugin can test a local schema against a registered subject, which makes it a useful CI checkpoint. Confluent also recommends pre-registering schemas through a controlled delivery process in production instead of relying on producer auto-registration.
A simplified Maven step can be as small as:
mvn io.confluent:kafka-schema-registry-maven-plugin:test-compatibilitySee the Schema Registry Maven plugin documentation for the plugin configuration to add to your pom.xml, including the registry URL and the subjects to check.
The command itself is not especially interesting. What matters is that an incompatible change fails while the developer is still looking at the pull request, rather than during application deployment.
There is one more boundary: application startup
CI can catch plenty of mistakes, yet it is still working with assumptions made before the service reaches its target environment.
At startup, the service connects to the Schema Registry endpoint configured for that environment. This is where potential drift after CI becomes visible.
The startup check asks whether the subject exists, whether its effective compatibility type is the one the service expects, whether the local schema is compatible with the latest registered version, and what to do if Schema Registry cannot be reached.
Those questions matter most when the service is pointed at the Schema Registry endpoint it will actually use.
The two checks happen at different moments for a reason. CI deals with the proposed change before it is promoted. Startup validation deals with the registry state the deployed service actually sees. Some overlap is useful because the failure modes are not identical.
The flow is simple:
One concrete implementation of the pattern
To address this startup gap, I built a small Spring Boot starter that applies the same contract check at startup. In Spring Boot, a starter packages the dependencies and auto-configuration needed to add a capability to an application with minimal setup.
The starter accompanying this blog is focused. It leaves serialization and Schema Registry behavior alone and concentrates on one thing: checking the contract the application expects before the service becomes ready.
Add the starter dependency without changing application source code, then configure the registry and the contract the service expects:
<dependency>
<groupId>io.github.mathias82.spring.kafka</groupId>
<artifactId>spring-kafka-contract-starter</artifactId>
<version>0.2.4</version>
</dependency>kafka:
contract:
enabled: true
compatibility: BACKWARD
registry:
url: ${SCHEMA_REGISTRY_URL}
subjects:
- name: order-events-value
schema-file: classpath:schemas/order-event.avsc
schema-type: AVROThat is enough for the startup check to run before the service becomes ready.
For each configured subject, the starter checks the assumptions declared by the application. It verifies that the subject has a registered version, resolves the effective compatibility type, compares it with the expected type, and asks Schema Registry whether the local schema is compatible with the latest registered version.
If any of those assumptions is false, startup can fail before the service becomes ready to receive traffic. For an incompatible local schema, the failure is explicit in the application logs:
IncompatibleSchemaException: Schema is NOT compatible for subject: order-events-valueIf the service already knows enough to reject the contract at startup, I would rather stop there than discover the problem on the first send or after a consumer has started processing traffic.
The starter also makes the behavior explicit when Schema Registry is unavailable, so teams can choose whether startup should stop or continue with a warning.
A runnable example
Along with the starter, I created a companion repository to simulate a client application encountering failure scenarios.
The demo runs Kafka and Confluent Schema Registry in Docker and consumes the published starter artifact rather than a local library build. That keeps the example close to the way another developer would actually use it.
The demo keeps the important cases visible: a normal end-to-end flow, a compatible change, a breaking change rejected at startup, and the service behavior when Schema Registry is temporarily unavailable. The repository contains the full commands and walkthrough.
Consider a service built for OrderEvent v1: it expects orderId (string), amount (double), and createdAt (string with a default). Another team proposes v3, which keeps orderId, removes amount and createdAt, and adds status (string without a default). Under BACKWARD compatibility, Schema Registry checks whether v3 can read data written with v1. Dropping fields is allowed, but the new required status field has no value in v1 records, so Schema Registry would reject the proposed registration. (Dropping amount, which has no default, would also break existing v1 consumers, but that is a FORWARD concern that BACKWARD alone does not guard against.) The runnable demo tests the same incompatible schemas at startup: v3 is supplied as the application's local schema against registered v1, and the starter rejects it.
The demo uses Spring Boot because it is my preferred application framework. The pattern itself is broader: validate compatibility before promotion, then assert the deployed contract again before the application becomes ready. The same approach can be implemented in another JVM framework or in a service written in a different language.
What compatibility checks still do not cover
These checks only cover schema compatibility. They do not tell us whether a consumer interprets an optional field correctly or whether the business meaning of a field changed. For higher-risk events I still test older payloads, producer serialization, and the consumer paths that matter.
Treat schema changes like API changes
Kafka decouples producers and consumers at runtime. But the meaning of the data still ties them together.
Schema changes deserve much the same treatment as API changes: keep them in version control, review them, test compatibility, and exercise the important producer and consumer paths. The earlier an incompatible schema change is visible, the easier it is to fix.
Catch the contract break before the traffic does.
Until the next commit.
- Schema Registry compatibility types↩