SAP error handling in Integration Suite: patterns and best practices
Error handling is the difference between an integration estate you can trust and one that fails silently. This pillar explains the SAP Cloud Integration building blocks: the exception subprocess, message statuses (RETRY, ESCALATED, FAILED), the retry patterns, idempotency and dead-letter handling, and the alerting that ties it together. It links to the detailed articles on the exception subprocess, retry mechanisms and monitoring.

Good error handling in SAP Integration Suite means every integration flow decides, by design, what happens when processing fails. You use an exception subprocess to classify the error, choose Error End or Escalation End to set the message status, add retry for recoverable errors through JMS or a Data Store, and alert on failures so nothing fails silently.
Why does error handling deserve its own discipline?
Integrations fail for ordinary reasons: a target system is briefly down, a certificate expires, a payload is malformed, a mapping meets data it did not expect. The question is not whether errors happen but what your integration does when they do. A flow with no error handling either fails silently or returns a confusing fault to the caller, and the business finds out when a process breaks two days later.
Error handling is therefore a design discipline, not an afterthought you bolt on. This guide is the reliability hub: it explains the SAP Cloud Integration building blocks and links to the deep-dive articles. If you are new to the platform, start with the complete guide to SAP Integration Suite.
What message statuses does Cloud Integration use?
Quick answer: COMPLETED, RETRY, ESCALATED and FAILED, and your design chooses which one a failure ends in.
Understanding the statuses is the foundation, because the status is the signal operations monitors and the outcome your exception handling produces.
| Status | Meaning | How it happens |
|---|---|---|
| COMPLETED | The message processed successfully | Normal end |
| RETRY | An error occurred and automatic retry has started | Recoverable error in an asynchronous (JMS or Data Store) scenario |
| ESCALATED | A handled error, flagged for monitoring | Exception subprocess ends with an Escalation End event |
| FAILED | Processing failed with no more retries possible | Exception subprocess ends with an Error End event, or retries are exhausted |
The practical implication: if you want automatic retry, the error must stay unhandled enough for the runtime to retry it; if you want to stop and classify, choose Error End for a hard failure or Escalation End for a monitored one.
What is the exception subprocess?
The exception subprocess is the catch block of an integration flow. When an error is raised anywhere in the main process, control passes to the exception subprocess, where you decide the outcome. You can log the error, build a meaningful error response, store the payload for later, or end with a specific status.
Because nearly every robust iFlow needs one, it is worth building a standard template with consistent logging, a correlation ID and a clear ending. The step-by-step is in SAP CPI exception subprocess: a practical guide.
How do I decide recoverable versus permanent errors?
Quick answer: recoverable errors are transient and should be retried; permanent errors are deterministic and should fail fast with an alert.
This single decision drives the design:
- Recoverable (transient): connection timeouts, a target temporarily unavailable, throttling, a transient lock. The right response is retry with backoff.
- Permanent (deterministic): schema validation failure, a business rule rejection, a missing mandatory field, an authorisation error. Retrying will fail again, so fail fast, capture the payload, and alert a human.
Mixing these up is the most common error-handling mistake. Retrying a permanent error wastes resources and hides the real problem; failing a transient error creates unnecessary incidents and wakes people up at night for something that would have fixed itself.
How does retry actually work?
For asynchronous scenarios, you decouple receipt from delivery using a store, then reprocess from that store:
- A JMS queue persists the message and the adapter reprocesses it automatically, tracking attempts with the SAPJMSRetries header.
- A Data Store persists the payload and a second iFlow or a looping call reprocesses it on a schedule, tracking attempts with SAP_DataStoreRetries.
The choice between them has real trade-offs in throughput, ordering and operational visibility. The full comparison, with a decision table, is in SAP CPI retry mechanism: JMS vs Data Store.
What about backoff and retry limits?
Retry is not just repeat immediately. Two parameters matter:
- Backoff: wait progressively longer between attempts so you do not hammer a struggling target. A target that is down for thirty seconds should not receive hundreds of calls in that window.
- Maximum attempts: cap the number of retries. Infinite retry is an anti-pattern that hides real problems and can turn one failing target into a platform-wide backlog.
Set these per interface based on how the target behaves, not with a single global number.
What about dead-letter handling?
Retry cannot be infinite. After the maximum attempts, a message that still fails should move to a dead-letter path: a queue or store where it waits for human attention, with enough context (original payload, headers, error, correlation ID) to diagnose it. A dead-letter path with no alert is just a hidden backlog, so always pair it with monitoring.
How do I know something failed?
Error handling that nobody sees is not error handling. Alert on FAILED messages, on growing RETRY backlogs, on escalations, and on dead-letter arrivals. SAP provides alerting through the platform, SAP Alert Notification service for pushing alerts to people and tools, and SAP Cloud ALM for integration and exception monitoring across SAP cloud and on-premise components. The practical setup is in how to monitor SAP Integration Suite.
How do I give errors business meaning?
A technical stack trace is not useful to the person who has to act. Good error handling attaches business context: which interface, which business process, which document or key, and a correlation ID that ties the failure to the wider transaction. When an alert carries that context, triage takes minutes instead of hours. This is why logging in the exception subprocess should be structured and consistent, not ad hoc.
What does a good standard look like?
Standardise error handling so every team does it the same way:
- A reusable exception-subprocess template with consistent logging and correlation IDs.
- A documented rule for recoverable versus permanent errors.
- A retry policy with backoff and a maximum attempt count.
- A dead-letter destination and a rule for how long payloads are kept.
- Alerts wired to the team that owns the interface, carrying business context.
What are the anti-patterns to avoid?
- Catch and ignore: swallowing an error so the message looks successful. This is the worst case, because the business data is simply lost.
- Retry everything forever: including permanent errors, which never succeed.
- One generic alert channel nobody owns, so alerts become noise.
- No correlation ID, so you cannot trace a failure across steps and systems.
How do I correlate errors across a business process?
Quick answer: attach a correlation ID at the entry point and carry it through every step, log and alert.
A single business transaction often crosses several iFlows and systems. Without a shared identifier, a failure in one hop is impossible to connect to the originating request. Generate or capture a correlation ID when a message first enters the platform, propagate it in headers across every step and call, and log it in every exception subprocess. When an alert fires, that ID lets an operator reconstruct the whole journey in minutes. This is also what makes monitoring and AI-assisted triage actually useful, because the signal is connected rather than fragmented.
What belongs in an error-handling runbook?
Error handling is not only design; it is operations. For each production interface, write a short runbook covering:
| Section | What it answers |
|---|---|
| What it does | The business process and systems involved |
| Common failures | The errors seen in practice and whether they are recoverable |
| Retry and dead-letter | The retry policy and where exhausted messages go |
| Who to call | The owning team and escalation path |
| How to replay | The safe way to reprocess a dead-lettered message |
A runbook turns a 2 a.m. incident from an investigation into a procedure. It also makes the interface auditable, which matters for regulated estates.
How do I handle partial failures and compensation?
Some processes touch several systems, and a failure halfway through leaves them inconsistent. Prefer idempotent, retriable steps so a replay simply completes the work. Where a true rollback is needed, design an explicit compensation step rather than hoping the systems sort themselves out. The honest guidance is to avoid distributed transactions where you can and design for eventual consistency with clear reconciliation, because integration rarely gives you a clean two-phase commit across heterogeneous systems.
How does error handling relate to testing?
You cannot claim an interface handles errors well if you never test the error paths. Include negative cases in your regression pack: a target that times out, a malformed payload, an authorisation failure. Confirm that recoverable errors retry and that permanent errors fail cleanly with a useful alert. Testing the unhappy path is where real reliability is proven, and it is covered in the gated guide on testing SAP integrations.
Which errors are usually recoverable, and which are permanent?
Quick answer: transport and availability failures are usually recoverable and should be retried; data, mapping and validation failures are usually permanent and should go straight to a human or a business owner. Authentication failures sit in between: retrying will not fix them, but a configuration fix will.
| Failure | Typical cause | Recoverable? | Right response |
|---|---|---|---|
| HTTP 503, 504, timeout, connection refused | Target down or overloaded | Yes | Retry with backoff |
| HTTP 429 | Target throttling the caller | Yes | Retry with longer backoff; review volume |
| HTTP 400 or 422, mapping error | Bad or unexpected data | No | Stop, alert the data owner |
| HTTP 401 or 403 | Expired credential, missing role | Not by retry | Alert integration team; fix configuration |
| HTTP 404 | Wrong endpoint or missing record | Usually no | Investigate; route by business meaning |
| HTTP 409 | Duplicate or conflicting update | Depends | Treat as success if idempotent duplicate |
Treat this table as a starting classification, not a law. The point is to decide in advance, per interface, so that the exception subprocess can route automatically instead of every failure landing in the same undifferentiated pile.
How do I read a failed message in the monitor?
Quick answer: open the message in Monitor Message Processing, read the status and error details first, then use the message processing log to see which step failed. Raise the log level only for as long as you need it.
The message processing log records each step the message passed through, so the last successful step tells you where to look. The error details usually contain the adapter's exception text, including the HTTP status and often the target's response body. If that is not enough, temporarily raise the log level for the integration flow. Trace captures payloads and is deliberately time-limited by SAP, which is a reminder that payload capture has data-protection consequences. A good runbook tells the on-call engineer exactly which of these to check, in what order, for each critical interface.
How do I capture HTTP and OData error details for routing?
Quick answer: inside the exception subprocess, read the caught exception, extract the status code and response body, and store them as exchange properties that a router can branch on.
The exchange property CamelExceptionCaught holds the exception object, and the simple expression for exception.message gives you its text. For HTTP receiver failures, the exception exposes the response status code and body, which a short Groovy script can copy into properties such as an error code and an error text. A router step then sends 5xx and timeouts to the retry path, 4xx to the business-error path, and anything unrecognised to a safe default that fails loudly. This one pattern is the bridge between "something failed" and "the right thing happened next". The exception subprocess guide walks through the build step by step.
What does error-handling maturity look like?
Quick answer: most teams progress from reactive firefighting, through standardised handling and owned alerting, towards automated, self-correcting operations. Knowing your level tells you what to invest in next.
| Level | What it looks like | Next investment |
|---|---|---|
| 1. Reactive | Users report failures; fixes are ad hoc | Basic alerting on failed messages |
| 2. Standardised | Shared exception subprocess, retry on async flows | Classification and business-meaningful errors |
| 3. Owned | Every critical interface has an owner, runbook and SLA | End-to-end correlation and process monitoring |
| 4. Autonomous | Known failures self-resolve; humans handle exceptions | Continuous improvement from incident data |
Level four is not about removing people; it is about reserving their attention for problems that genuinely need judgement. Our gated guide on the autonomous integration operations model describes that end state in detail.
A scenario: a supplier invoice that fails at 2 a.m.
Quick answer: good error handling turns a silent overnight failure into an automatic retry, a precise alert, and a fix the business never notices.
A supplier invoice arrives at 2 a.m. and the S/4HANA posting call times out. Because the flow is asynchronous and backed by a JMS queue, the message returns to the queue and retries with exponential backoff. S/4HANA recovers at 2.20 a.m. and the retry succeeds; nobody is woken. The next night, a different invoice fails with a 422 because the supplier sent an unknown tax code. That is permanent, so the exception subprocess classifies it as a business error, ends the message as failed with a clear reason, and notifies the accounts-payable data owner rather than the integration on-call. By 9 a.m. the tax code is corrected and the message is reprocessed. Two failures, two very different responses, both decided by design rather than by whoever happened to notice.
When should I not add custom error handling?
Quick answer: do not wrap every step in bespoke handling, and never use an exception subprocess to hide failures. Use platform retry where it suffices and keep custom logic for genuine business decisions.
Over-engineering error handling is a real risk. If an asynchronous flow only needs to survive brief outages, the JMS sender's built-in retry with backoff may be all it needs. Ending every exception with a normal end event makes the monitor show success while data is lost, which is worse than having no handling at all. Simplicity, consistency and honesty about failure beat cleverness every time.
Where governance meets reliability
Error-handling standards belong in your change control. A change that weakens retry, removes an alert, or turns an Error End into a silent catch is a control change, not a cosmetic one, and should be reviewed as such. For the governance angle, the gated playbook on SAP integration error handling goes deeper into the operating model, and the integration governance playbook covers how this fits day-to-day operations. For the platform overview, return to the SAP Integration Suite complete guide.
Key takeaways
- Every integration flow should handle exceptions explicitly rather than letting them fail silently.
- The exception subprocess classifies errors: an Error End event marks the message FAILED, an Escalation End event marks it ESCALATED.
- RETRY means automatic reprocessing has started; FAILED means no more retries are possible.
- Distinguish recoverable errors (retry) from permanent errors (fail fast and alert).
- Idempotency and a dead-letter path are prerequisites for safe retry.
- Standardise one error-handling template and wire alerts to the team that owns each interface.
Questions
What is the exception subprocess in SAP CPI?
It is a block inside an integration flow that catches processing errors. You decide what it does: log, transform an error response, trigger a retry, or end the message with an Error End or Escalation End event that sets the final message status.
What is the difference between FAILED, ESCALATED and RETRY?
RETRY means an error occurred and automatic retry has started. FAILED means processing ultimately failed with no more retries possible, which an Error End event produces. ESCALATED is produced by an Escalation End event and is used to flag a handled error for monitoring.
How do I make an error recoverable?
Keep the message in an asynchronous store (a JMS queue or a Data Store) and let the runtime reprocess it. For JMS-based retry, do not swallow the exception in a way that finalises the message, or the retry will not happen.
Why does idempotency matter for error handling?
Retry means the same message can be processed more than once. If the receiver is not idempotent, a retry can create duplicates. Design receivers and flows so repeat processing is harmless.
Where can I see failed messages?
In the message monitor in Cloud Integration, which shows per-message status, logs and attempts, and in SAP Cloud ALM for integration and exception monitoring across the landscape. Alert on FAILED and escalations so you do not rely on checking manually.
Related reading
See what Spanovix would fix in your landscape
Bring your hardest interfaces. In a short working session we show where the agents cut failures, manual work and risk for your teams.
