AEM Event Handling — Observing Repository Changes Without Building Fragile Listeners
A practical developer and architect guide to event handling in AEM, covering JCR Observation, Sling ResourceChangeListener, OSGi EventHandler, event boundaries, asynchronous handoff, filtering, and production-safe listener design.
AEM Event Handling — Observing Repository Changes Without Building Fragile Listeners
Some AEM requirements begin with a repository change rather than an HTTP request or a scheduled time.
An asset is uploaded.
A content fragment changes.
A page is created.
A property is updated.
A workflow changes repository state.
Something else in the application needs to react.
The first implementation idea is often:
Add a listener and run the logic when the node changes.
That works for small examples.
In production, event handling becomes more complicated because one business action can generate several repository changes, listeners can execute on more than one runtime instance, the callback may be invoked frequently, and the work triggered by the event may be much heavier than the listener should perform directly.
So I do not start event-driven AEM design by asking:
Which listener API should I use?
I start with:
What happened, which layer should observe it, and how much work is safe to perform at that boundary?
Event Handling Is a Trigger Boundary
An event listener should normally detect something and hand responsibility to the appropriate service.
It should not become the entire feature.
For example:
Asset change
Listener detects relevant change
Business service decides what it means
Background work runs if needed
The listener owns detection.
The service owns business behavior.
If the resulting work is expensive, retryable, or needs to survive runtime interruption, the listener can enqueue a Sling Job instead of performing all work synchronously.
That boundary connects directly to Chapter 29.
Events tell us something happened.
Jobs give us a stronger model for work that now needs processing.
The Main Event Mechanisms You Will See in AEM
Several APIs can appear in AEM event-driven implementations.
The most common ones are:
- JCR Observation
- Sling
ResourceChangeListener - OSGi
EventHandler - Sling Jobs for asynchronous processing after detection
These mechanisms operate at different abstraction levels.
Choosing between them should follow the information the requirement actually needs.
JCR Observation
JCR Observation works at the repository level.
A listener can observe events such as node creation, removal, movement, and property changes.
Conceptually:
ObservationManager observationManager =
session
.getWorkspace()
.getObservationManager();
observationManager.addEventListener(
eventListener,
Event.NODE_ADDED
| Event.PROPERTY_CHANGED,
"/content/dam/myproject",
true,
null,
null,
false
);
This is a low-level repository mechanism.
It can be appropriate when the implementation genuinely needs JCR-level event information.
But most application code does not need to begin this low in the stack.
AEM application code usually works with Sling resources rather than directly with JCR nodes.
If the requirement is:
React when resources below this content path are added, changed, or removed.
then Sling's resource-level listener is often the cleaner boundary.
Sling ResourceChangeListener
ResourceChangeListener observes resource changes using the Sling
Resource API abstraction.
A simple listener can look like this:
@Component(
service = ResourceChangeListener.class,
property = {
ResourceChangeListener.PATHS
+ "=/content/dam/myproject",
ResourceChangeListener.CHANGES
+ "=ADDED",
ResourceChangeListener.CHANGES
+ "=CHANGED",
ResourceChangeListener.CHANGES
+ "=REMOVED"
}
)
public class AssetResourceChangeListener
implements ResourceChangeListener {
@Override
public void onChange(
List<ResourceChange> changes) {
for (ResourceChange change
: changes) {
// Inspect the change and hand off
// relevant work.
}
}
}
The listener receives a list because several changes can be delivered together.
That alone is a useful reminder not to design the callback around the assumption that exactly one business action equals exactly one callback.
Filter as Close to the Event Boundary as Possible
A broad listener creates unnecessary work.
This:
/content
is very different from:
/content/dam/myproject
If the requirement concerns one project DAM area, listening to the entire content tree means the application receives changes it will immediately discard.
Filtering should begin with the listener registration:
- Relevant path
- Relevant change type
- Relevant resource provider behavior where applicable
Then application-level checks can narrow the event further.
For example, the listener may care only about assets under:
/content/dam/myproject/products
and only when a particular metadata property or processing state makes the asset relevant.
The earlier irrelevant events are removed, the less unnecessary work reaches the application.
ADDED, CHANGED, and REMOVED Are Repository/Resource Events
An ADDED event does not automatically mean:
A business user created one complete asset and it is ready for my integration.
A single AEM operation can create or modify several resources.
Asset ingestion is a good example.
Creating an asset can result in changes to the asset resource, metadata, renditions, processing state, and related repository structures.
A listener that reacts to every low-level change as though it were a complete business event can trigger the same downstream work several times.
So event type is only the first filter.
The application still needs to decide whether the current repository state represents the business condition it cares about.
Do Not Assume Event Delivery Equals Business Exactly-Once
This is one of the most important boundaries in event-driven code.
A listener notification tells us that a change was observed.
It does not prove that a business operation should execute exactly once.
The same logical content operation may generate multiple relevant changes.
The application may restart.
Cluster behavior may affect where notifications are observed.
The downstream work may itself fail and retry.
If duplicate business processing is harmful, the application needs its own idempotency or state check.
The event mechanism is not a business deduplication engine.
Keep Heavy Work Out of the Listener
Consider this implementation:
@Override
public void onChange(
List<ResourceChange> changes) {
for (ResourceChange change
: changes) {
externalProductApi
.synchronize(
change.getPath()
);
}
}
The listener now owns an external network call.
If that API becomes slow, the event callback becomes slow.
If it fails temporarily, the listener needs retry logic.
If several hundred changes arrive, the callback can spend significant time performing integration work.
That is too much responsibility for the event boundary.
A stronger design is:
@Override
public void onChange(
List<ResourceChange> changes) {
for (ResourceChange change
: changes) {
if (!isRelevant(change)) {
continue;
}
queueProductSync(
change.getPath()
);
}
}
The listener performs lightweight filtering and creates background work.
A Sling Job then owns the retryable processing.
Event to Job Is a Useful Production Pattern
Suppose product content changes under:
/content/dam/myproject/products
and an external search index must be refreshed.
The listener can create a job:
private void queueProductSync(
String resourcePath) {
Map<String, Object> properties =
new HashMap<>();
properties.put(
"resourcePath",
resourcePath
);
jobManager.addJob(
"myproject/product/reindex",
properties
);
}
The JobConsumer can then:
- Load the current resource state.
- Validate that the content still exists and is relevant.
- Perform the integration.
- Return the appropriate job result.
- Retry temporary failures according to queue policy.
The listener does not need to own any of those concerns.
Why the Consumer Should Re-read Current State
An event describes a change that already happened.
By the time an asynchronous job runs, the resource may have changed again.
Suppose metadata changes three times quickly:
Version A
Version B
Version C
Three listener notifications may produce work.
If the actual requirement is:
Make the search index reflect the latest product state.
then serializing the entire old resource state into each job payload is usually the wrong model.
A smaller payload such as:
resourcePath=/content/dam/myproject/products/product-a
lets the consumer read the current state when it processes the job.
Whether every intermediate state matters is a business decision.
For many synchronization use cases, only the latest state matters.
Removed Resources Need Different Handling
Re-reading current state works for additions and updates.
Removal is different.
After:
ResourceChange.ChangeType.REMOVED
the resource may no longer exist.
If downstream cleanup needs information that existed before deletion, the path alone may not be enough.
That requirement should be designed explicitly.
For example, an external index may need only the deleted content identifier or path.
If that identifier can be derived from the event path, the job can use it.
If cleanup needs metadata that disappears with the resource, the application needs a different strategy for preserving the information required by the deletion operation.
Do not assume a delete event can later reconstruct the deleted resource.
OSGi EventHandler
OSGi Event Admin is another event mechanism commonly seen in AEM.
An EventHandler subscribes to one or more event topics.
For example:
@Component(
service = EventHandler.class,
property = {
EventConstants.EVENT_TOPIC
+ "=myproject/content/published"
}
)
public class ContentPublishedEventHandler
implements EventHandler {
@Override
public void handleEvent(
Event event) {
String path =
(String) event.getProperty(
"path"
);
// Delegate relevant behavior.
}
}
This is different from directly observing repository resources.
The contract is an event topic and its properties.
That makes Event Admin useful when one component needs to publish an application event and another component needs to react without a direct Java service call between them.
Repository Event vs Application Event
This distinction helps prevent overusing repository listeners.
Suppose a business service knows that a product was successfully approved.
It could update repository state and expect another component to infer the business meaning from a low-level property change.
Or it could explicitly publish an application event such as:
myproject/product/approved
with the required identifier.
The second design can make the business meaning much clearer.
Repository listeners are useful when the repository change itself is the trigger we need to observe.
Application events are useful when the application already knows the semantic event that occurred.
Do not force every business event to be rediscovered from JCR changes.
Event Topics Are Contracts
An event topic such as:
myproject/product/approved
connects the producer and consumer.
The properties form part of that contract too.
For example:
Dictionary<String, Object> properties =
new Hashtable<>();
properties.put(
"productId",
productId
);
Event event =
new Event(
"myproject/product/approved",
properties
);
eventAdmin.postEvent(event);
Consumers now depend on:
- The topic name
- The meaning of the event
- The expected properties
- The type of each property
Treating these as arbitrary strings spread across classes makes event-driven code difficult to change.
Constants or a small event-contract abstraction can keep producers and consumers aligned without turning the event bus into a generic dumping ground.
sendEvent vs postEvent
OSGi Event Admin supports two delivery styles:
sendEvent(...)performs synchronous delivery and returns after delivery completes.postEvent(...)initiates asynchronous delivery and returns before delivery completes.
That is a delivery choice, not a persistence guarantee. If the resulting work must survive interruption, retry, or runtime replacement, hand the work to Sling Jobs rather than treating Event Admin as a durable queue.
Do Not Turn Event Admin Into a Job Queue
An event communicates that something happened. A job represents work that needs reliable processing.
If an application event requires retryable downstream work, keep the
EventHandler lightweight and enqueue a Sling Job. Event Admin and
Sling Jobs solve different responsibilities.
Listener Loops
Event-driven repository code can accidentally trigger itself.
Imagine:
- Listener observes a metadata change.
- Listener writes
processingStatus=complete. - That write generates another relevant change.
- Listener reacts again.
- It writes again.
Even when the loop eventually stops, unnecessary repeated processing can occur.
The first protection is usually a precise business condition.
For example, do not react merely because something under the asset changed.
React only when the current state requires the downstream action.
If the listener itself writes repository state, make sure its own write cannot continuously satisfy the trigger condition.
Event Storms
A bulk content operation can generate a large number of changes.
Examples include:
- Asset migration
- Metadata updates
- Package installation
- Bulk content import
- Workflow-driven updates
- Automated maintenance
A listener that is harmless when one author edits one page may behave very differently when 50,000 resources change.
Before approving a listener design, I want to know:
- How broad is the observed path?
- How many changes can one business operation generate?
- Does the listener perform repository reads for every event?
- Does it create one job for every low-level change?
- Can equivalent work be collapsed?
- Can the downstream system handle the resulting rate?
Event-driven design needs a volume model, not just a functional test.
Coalescing Repeated Work
Suppose an asset receives ten metadata changes in a few seconds.
If the downstream requirement is simply:
Reindex the latest asset state.
creating ten expensive indexing operations is wasteful.
Possible strategies include:
- Detect whether equivalent work is already pending.
- Store a durable "needs reindex" state and let processing clear it.
- Use a reconciliation job that collapses repeated changes.
- Make repeated indexing cheap and idempotent if volume is low.
The right strategy depends on scale and business semantics.
The listener API itself does not solve coalescing.
Repository Access Inside Listeners
A listener may need repository access to inspect current state.
The same resolver rules from Chapters 26 and 27 apply.
Do not store a ResourceResolver as long-lived listener state.
Do not assume a request resolver exists.
If a service resolver is required, create it for the operation, use the appropriate subservice, and close it in the same execution scope.
If the listener only needs to enqueue the changed path, it may not need repository access at all.
Avoid opening a resolver just because the event is repository-related.
Event Handler Thread Safety
Listener and handler components are OSGi services.
Their methods may be invoked over the lifetime of the component and potentially under concurrent workloads.
Mutable instance fields such as:
private String currentPath;
should not be used to hold per-event state.
Keep event-specific data in local variables.
Shared collaborators should themselves be designed for the concurrency model in which they are used.
The same service thread-safety principles from Chapter 25 apply here.
A Practical Asset Metadata Example
Assume the requirement is:
When product asset metadata changes, update the external product search index.
A weak design is:
onChange(...)
-> open resolver
-> read asset
-> call external API
-> retry manually
-> update repository
all inside the listener.
A stronger design separates the responsibilities.
Listener
Detects relevant changes below the product asset path and queues work.
Job
Represents the reindex request and provides retry/recovery semantics.
ProductIndexService
Loads the current product state and communicates with the external index.
Repository service identity
Provides only the repository access required by the indexing operation.
The event path remains lightweight even when the external integration becomes slow or temporarily unavailable.
Production Failure: The Listener Fires Too Often
The first thing I check is whether the implementation is confusing low-level resource changes with business events.
Questions include:
- Is the path too broad?
- Are several child resources changing for one operation?
- Is the listener reacting to its own writes?
- Are external changes also being observed?
- Is every
CHANGEDevent really relevant? - Is equivalent work being queued repeatedly?
Adding more logging inside the downstream integration does not solve an event-selection problem.
Start at the listener boundary.
Production Failure: Local Works, Cloud Produces Duplicate Work
Local development usually has a much simpler runtime topology.
When duplicate processing appears after deployment, inspect:
- Whether events are local or external.
- Whether multiple instances can observe or react to the same logical change.
- Whether the listener queues duplicate logical work.
- Whether the job itself can retry.
- Whether the business operation is idempotent.
Do not solve this only by trying to force the listener onto one JVM.
The stronger design makes duplicate processing safe or prevents duplicate business work at the correct boundary.
Production Failure: Event Handling Slows Content Operations
If authoring or repository operations become slow after introducing a listener, inspect what the callback is doing.
A listener should not routinely perform:
- Slow network calls
- Large repository traversals
- Expensive transformations
- Long retry loops
- Bulk writes unrelated to the immediate detection step
Move expensive work behind an asynchronous processing boundary.
Also verify that the listener is not receiving a much larger event set than expected.
Production Failure: A Delete Cannot Be Processed
This usually happens when the implementation reacts to REMOVED and
then tries to load the deleted resource.
By that point:
resolver.getResource(
deletedPath
);
can correctly return null.
The fix is architectural.
Decide what information the downstream deletion requires and make sure that information survives long enough to process the removal.
Do not treat a missing deleted resource as a repository failure.
A Troubleshooting Sequence
When an event-driven feature behaves incorrectly, I trace it from the source rather than starting at the downstream integration:
- Confirm that the expected repository or application event occurred.
- Verify listener path/topic/change-type registration.
- Check business filtering and whether several low-level events created duplicate logical work.
- Confirm that the listener handed off expensive work instead of blocking the callback.
- Inspect downstream job retry/idempotency behavior.
- Verify the repository identity used by background processing.
- For deletes, confirm the design does not expect the removed resource to still exist.
This keeps the investigation close to the event boundary before blaming the final integration.
Architect Perspective
Event handling introduces a time gap between a change and the reaction to it. The architecture therefore needs to define the event contract, duplicate behavior, work durability, current-state validation, delete semantics, and repository identity.
If the listener itself contains most of the feature, those boundaries are usually too tightly coupled.
Local and External Resource Changes in AEM as a Cloud Service
Cloud topology changes how I think about observation code.
AEM as a Cloud Service always runs in a cluster. Adobe also calls out observation specifically: JCR and Sling resource events cannot be treated as guaranteed local execution because an instance can disappear and another active instance may observe the change as an external event.
That means this design is unsafe:
Listen only for a local change, perform critical business work immediately, and assume the callback will always happen on the instance where the repository change originated.
For a ResourceChangeListener, locality should be an explicit decision.
A listener that implements only:
ResourceChangeListener
receives the resource changes appropriate to that registration.
If the implementation also needs changes marked as external by the resource provider, it can implement:
ExternalResourceChangeListener
as well.
For example:
@Component(
service = ResourceChangeListener.class,
property = {
ResourceChangeListener.PATHS
+ "=/content/dam/myproject",
ResourceChangeListener.CHANGES
+ "=CHANGED"
}
)
public class ProductAssetChangeListener
implements ResourceChangeListener,
ExternalResourceChangeListener {
@Override
public void onChange(
List<ResourceChange> changes) {
// Keep the callback lightweight.
}
}
I would not add ExternalResourceChangeListener automatically.
The application first needs to decide whether externally reported changes should create the same business work as local changes.
That decision belongs in the architecture because accepting both local and external changes can affect duplicate-work behavior.
Observation Is Not a Guaranteed Work Queue
This is worth stating directly for AEM as a Cloud Service.
Adobe's current development guidance says observation events must be used with care because execution cannot be guaranteed locally. An instance can be stopped while asynchronous work is happening, and topology changes can affect which instance observes the event.
So an observation callback should not be treated as the only durable record that critical work exists.
If the requirement is:
Whenever this change happens, this integration must eventually complete.
then the design needs a recoverable state or durable processing mechanism beyond a transient listener callback.
Depending on the use case, that may be:
- A Sling Job created from the observed change.
- A durable repository state that a reconciliation process can rediscover.
- A business record that indicates downstream processing is still pending.
- An external event/integration mechanism designed for that delivery requirement.
The important part is not to confuse change notification with guaranteed completion of business work.
Reconciliation Makes Event-Driven Processing Safer
An event-driven path is fast because it reacts close to the change.
A reconciliation path is useful because it can recover work that the event path did not complete.
Consider product indexing.
The event listener may immediately enqueue reindex work when product content changes.
Separately, the application can maintain enough durable state to answer:
Which products still need indexing?
A periodic reconciliation can then detect work that is still pending.
This gives the design two different strengths:
- Event handling provides low-latency reaction.
- Reconciliation provides recovery.
Not every listener needs a reconciliation process.
For business-critical integrations, however, it is worth deciding explicitly what recovers the system if one event is never turned into completed downstream work.
ResourceChangeListener Registration Should Express Intent
Listener registration is part of the design, not just annotation syntax.
For example:
@Component(
service = ResourceChangeListener.class,
property = {
ResourceChangeListener.PATHS
+ "=/content/dam/myproject/products",
ResourceChangeListener.CHANGES
+ "=ADDED",
ResourceChangeListener.CHANGES
+ "=CHANGED"
}
)
public class ProductAssetListener
implements ResourceChangeListener {
}
This tells us two useful things immediately:
- Which resource subtree matters.
- Which change types matter.
That is much easier to reason about than registering broadly and hiding all filtering inside:
if (...)
statements.
If several unrelated content areas require different reactions, separate listeners can also be clearer than one global listener containing a long chain of path checks.
The listener registration should make the expected event surface visible.
OSGi Event Filters Can Narrow Application Events
Topic filtering is not the only filtering available to an OSGi
EventHandler.
Event Admin supports an EVENT_FILTER service property using an OSGi
filter expression.
For example, if application events contain a property such as:
site=myproject
a handler can register with both the topic and a filter.
@Component(
service = EventHandler.class,
property = {
EventConstants.EVENT_TOPIC
+ "=myproject/content/changed",
EventConstants.EVENT_FILTER
+ "=(site=myproject)"
}
)
public class MyProjectContentEventHandler
implements EventHandler {
@Override
public void handleEvent(
Event event) {
// Only events matching the topic
// and filter reach this handler.
}
}
This can keep event selection at the registration boundary instead of subscribing broadly and discarding most events in Java.
I still keep business validation inside the handler or delegated service.
Registration filtering reduces noise.
It does not replace validation.
Avoid Event Chains That Hide the Business Flow
Event-driven code becomes difficult to troubleshoot when every handler publishes another event.
For example:
content/changed
product/reload
product/validated
product/index-requested
product/index-started
can look nicely decoupled on paper.
In production, it may become difficult to answer which component owns the actual operation and where the failure occurred.
I use application events when independent components genuinely need notification.
I do not replace normal service calls with events just to avoid direct dependencies.
If component A always needs component B to complete the next synchronous business step, a Java service dependency may be much clearer.
Event-driven boundaries are useful when the temporal decoupling is intentional.
Publishing an Event After Repository Persistence
Another subtle boundary appears when code both changes repository state and publishes an application event.
Consider:
resource.adaptTo(
ModifiableValueMap.class
).put(
"approvalStatus",
"approved"
);
eventAdmin.postEvent(
approvalEvent
);
resolver.commit();
The event can now be observed before the repository change has been successfully persisted.
A handler that loads the resource may see the old state.
If commit() fails, the application may even have announced a business
event for a state that never became durable.
The safer order for an event that represents committed repository state is usually:
resource.adaptTo(
ModifiableValueMap.class
).put(
"approvalStatus",
"approved"
);
resolver.commit();
eventAdmin.postEvent(
approvalEvent
);
Even this is not a distributed transaction between the repository and Event Admin.
The repository commit can succeed and event publication can still fail later.
If losing that notification is unacceptable, the application needs a durable handoff/reconciliation design rather than pretending the two operations are one atomic transaction.
Event Ordering Should Not Be Assumed Without a Business Contract
OSGi documents ordered delivery for events submitted through
postEvent(), but that does not guarantee that downstream business work
will complete in the same order once handlers delegate to parallel or
independently retried processing.
If business order matters, make it explicit. Depending on the workload, that may mean an ordered job queue, version checks, ignoring stale work, or serializing work for the same business identifier.
Do not infer business completion order merely because repository changes happened sequentially.
Stale Events Need a Current-State Check
Event-driven systems naturally create a gap between detection and processing.
Suppose the listener sees:
product status = approved
and queues work.
Before the job runs, an author changes the product back to:
product status = draft
If the consumer blindly executes the original "approved" action, it may publish or synchronize stale business state.
For current-state integrations, the consumer should load the latest resource and verify that the triggering condition is still true.
For historical/event-stream use cases, preserving the original state may be correct.
The requirement decides which model applies.
The listener should not make that choice accidentally.
Testing Event-Driven Code
I do not try to prove the entire event infrastructure through one large unit test.
The useful tests are around the decisions the application owns.
For a ResourceChangeListener, I would test:
- Relevant path is accepted.
- Irrelevant path is ignored.
- Relevant change type creates work.
- Unrelated change type does not.
- Several changes in one callback are handled correctly.
- Duplicate logical changes do not create unsafe duplicate business work.
- Removal does not assume the deleted resource still exists.
For an OSGi event handler:
- Expected topic/property contract is interpreted correctly.
- Invalid or missing event properties are rejected safely.
- Handler delegates to the correct service.
- Heavy processing is not embedded in the handler.
For the downstream job:
- Current repository state is revalidated.
- Temporary failure is retryable.
- Permanent invalid work is cancelled.
- Duplicate/retried execution is safe.
This keeps framework wiring tests small and puts most coverage around the business boundaries.
Integration Testing Matters More Than Local Listener Tests
A local SDK is useful for validating registration and basic behavior.
It does not reproduce every characteristic of the AEM as a Cloud Service topology.
Adobe explicitly requires cloud code to be cluster-aware, and its observation guidance calls out the possibility that a different active instance reacts to an event when topology changes.
For listener-heavy features, staging tests should therefore include realistic operations such as:
- Bulk content updates.
- Multiple related resource changes.
- Job handoff under load.
- Repeated updates to the same content.
- Deployment/restart scenarios where recovery matters.
- Author or Publish topology appropriate to the feature.
The purpose is not to prove that an annotation works.
It is to prove that the feature still reaches a correct business state when event delivery and processing are not as simple as one local JVM.
A More Complete Event-Driven Design
For a business-critical content integration, I prefer a design with explicit layers.
Detection
A narrowly registered ResourceChangeListener or application
EventHandler identifies the relevant occurrence.
Filtering
The handler decides whether the change actually represents work the application cares about.
Durable handoff
If the work needs retry or recovery, a Sling Job represents it.
Business processing
A reusable service loads current state and performs the integration.
Repository identity
Background repository access uses a scoped service identity.
Idempotency
Repeated event/job execution does not create invalid duplicate business effects.
Recovery
If completion is critical, durable state or reconciliation can rediscover unfinished work.
Observability
Logs correlate the resource/event identifier with the downstream job or business operation.
The listener remains the smallest part of the design.
That is usually a good sign.
Production Review Checklist
Before approving an event-driven implementation, I review these areas:
Area Question
Event source Are we observing the right abstraction: repository resource change or application event?
Registration Is the path/topic/change type as narrow as practical?
Cloud topology Do local/external events and instance replacement affect the design?
Work size Is the listener doing only lightweight detection/filtering?
Durability What happens if the callback does not lead to completed business work?
Duplicate work Can several low-level events represent one logical operation?
Ordering Does downstream completion need to preserve order?
Current state Should processing use event-time state or the latest repository state?
Delete handling Is enough information available after the resource disappears?
Repository access Which service identity performs background reads/writes?
Recovery Is reconciliation required for business-critical completion?
Observability Can we trace event → job → business operation during an incident?
This is more useful than reviewing only the listener class.
Most production failures happen at the boundaries around it.
Summary
AEM offers event mechanisms at different levels. JCR Observation works
at the repository level, ResourceChangeListener observes Sling
resource changes, and OSGi EventHandler handles topic-based
application events.
In AEM as a Cloud Service, observation code must be cluster-aware. Adobe explicitly warns that JCR and Sling observation processing cannot be assumed to execute locally because instances can be replaced and another active instance may react to the event.
Keep listeners narrow and lightweight. Delegate business behavior to services, use Sling Jobs when resulting work needs durable retry/recovery semantics, and use reconciliation when business-critical completion cannot depend on a single observation callback.
Most importantly, do not equate one low-level event with one exactly-once business operation.
What's Next
Chapter 31 — AEM Workflow Architecture
The next chapter moves deeper into workflows themselves: workflow models, steps, payloads, participant behavior, custom process steps, persistence, failure handling, and the boundary between workflow orchestration and reusable application services.
Enjoyed this chapter?
Get an email when I publish the next chapter. No spam — just new technical deep-dives.
Comments
Share feedback or questions about this blog post.
No comments yet. Be the first to share your thoughts.