Skip to content
Menu
Article

How to monitor SAP Integration Suite: alerting and failed messages

Monitoring turns an integration estate from something you hope works into something you know works. This article covers the message statuses to watch, the message monitor, how to alert on failed messages and retries, the difference between SAP Cloud ALM and SAP Alert Notification, and how to route alerts with ownership. It supports the reliability pillar.

How to monitor SAP Integration Suite: alerting and failed messages

Monitor SAP Integration Suite by watching message status in the message monitor, alerting on FAILED messages and growing RETRY backlogs, and using SAP Cloud ALM for landscape-wide integration and exception monitoring with SAP Alert Notification to push alerts to people and tools. Effective monitoring is proactive: you detect and act on failures before the business does, and route each alert to the team that owns the interface.

What does good monitoring actually mean?

Good monitoring means you find out that an interface failed before the business does, and you know which interface, why, and who should fix it. Weak monitoring means you learn about failures from an angry email two days later, then spend hours working out what broke. The difference is not more dashboards; it is watching the right signals and routing alerts to the right people with enough context to act.

Monitoring is the operational half of reliability. The design half is covered in SAP error handling in Integration Suite. The two only work together: monitoring is what makes your error handling visible.

Which signals matter most?

Quick answer: message status per interface, especially rising RETRY and FAILED counts.

Start with message status, because it is the most direct health signal for each flow.

StatusWhat it tells youWhat to do
COMPLETEDThe message processed successfullyBaseline; track volume trends for anomalies
RETRYA recoverable error occurred and reprocessing startedWatch the backlog; a growing RETRY count means a target is struggling
ESCALATEDA handled error flagged for monitoringReview; decide if it needs action
FAILEDProcessing failed, no more retriesAlert immediately; a human needs to act

A single FAILED message might be noise; a rising FAILED or RETRY count on one interface is a signal. Trend matters as much as the individual event, so monitor rates and backlogs, not just events.

How do I use the message monitor?

The message monitor is your per-message view. It shows processing status, logs, headers and attempt counts, and from it authorised users can inspect a failed message and, where the scenario supports it, reprocess it. When an alert fires, the monitor is where you go to understand the specific failure: which step failed, what the payload looked like, how many times it retried.

Use correlation IDs so you can follow one business transaction across steps and systems. If your exception subprocess logs a correlation ID on every error, triage becomes far faster, because you can jump straight from the alert to the exact message.

How do I set up alerting?

Alerting turns status into action. At a minimum, raise alerts on:

  • FAILED messages on any production interface.
  • Growing RETRY backlogs, which indicate a target is degraded even if nothing has fully failed yet.
  • Escalations, so handled-but-notable errors get reviewed.
  • Dead-letter arrivals, so exhausted-retry messages (see JMS vs Data Store) are never a silent backlog.

Cloud ALM vs Alert Notification: which do I use?

Quick answer: Cloud ALM detects and diagnoses; Alert Notification delivers the alert.

These solve different problems and are often used together:

ToolPurposeBest for
SAP Cloud ALMIntegration and exception monitoring across cloud and on-premise: message monitoring, search, correlation and alerting on failuresSeeing a failure in the context of the end-to-end process and diagnosing root cause
SAP Alert NotificationNotification delivery and subscription for BTP events and custom conditions, to channels like email, Slack or ticketingPushing an alert to the right humans or tools immediately

A common setup uses Cloud ALM to detect integration exceptions across the landscape and Alert Notification to fan those out to the teams that need to know.

How do I monitor across the whole landscape?

A single tenant view is not enough when a business process crosses S/4HANA, middleware and non-SAP systems. SAP Cloud ALM is designed for this: it provides integration monitoring and exception monitoring across components, with message search and correlation, so you can see a failure in the context of the end-to-end process rather than one iFlow at a time. For regulated estates, the gated playbook on monitoring SAP Integration Suite goes deeper into the operating model.

How should alerts be routed?

An alert that lands in a shared inbox nobody owns is noise. Route each alert to the team that owns the interface, and include enough context to act: the interface name, the business process it serves, the error, the correlation ID, and a link to the message in the monitor. Agree what each severity means and what the response time is. Monitoring without ownership does not reduce incidents; it just creates a louder backlog.

What are the monitoring anti-patterns?

  • Monitoring only platform availability, missing interface-level failures.
  • Alerting on every event rather than on trends, so alerts become noise and get ignored.
  • No correlation ID, so triage is slow.
  • No ownership, so alerts have no owner and no deadline.
  • No alert on dead-letter, so retry-exhausted messages pile up unseen.

What KPIs and SLAs should I track?

Quick answer: track per-interface success rate, failure rate, retry backlog, and time-to-resolution, against agreed targets.

Monitoring improves when it is measured. Useful indicators include:

KPIWhy it matters
Success rate per interfaceThe headline health signal; a drop is an early warning
Failure and escalation rateTrends reveal degrading targets before full outages
Retry backlog depthA rising backlog means a target is struggling
Time to detect and resolveMeasures how good your alerting and ownership actually are
Dead-letter volumeWork that needs human attention

Agree service levels per interface with the business, because a payment feed and an internal report do not deserve the same response time. KPIs without agreed targets are just numbers.

How do I keep alerts actionable and avoid fatigue?

The fastest way to make monitoring useless is to alert on everything. Alert fatigue sets in and real incidents get missed in the noise. Tune alerts to trends and thresholds rather than individual events, group related alerts, and make sure every alert has an owner and a clear action. If an alert does not require someone to do something, it should be a dashboard metric, not an alert. Review your alert rules periodically and delete the ones nobody acts on.

How does monitoring support an on-call rotation?

For critical integrations, someone needs to be reachable when things fail outside business hours. A workable setup routes high-severity alerts (FAILED on a critical interface, a growing dead-letter backlog) through SAP Alert Notification to an on-call channel or ticketing tool, with the interface runbook attached so the responder has context. Lower-severity signals go to a dashboard reviewed in business hours. Clear severities and ownership are what make on-call sustainable rather than exhausting.

How do I monitor an end-to-end business process?

A single iFlow view does not tell you whether a business process succeeded, because the process spans several interfaces and systems. SAP Cloud ALM's integration and exception monitoring, combined with consistent correlation IDs, lets you trace a transaction across hops and see where it stalled. This process-level view is what the business actually cares about: not whether an iFlow ran, but whether the order reached the warehouse. The gated guide on monitoring SAP Integration Suite goes deeper into building this capability.

How do I pull monitoring data into my own tools?

Quick answer: use the Cloud Integration OData API for message processing logs to query statuses, failures and timings programmatically, then feed them into your dashboards, ticketing or observability platform.

The message processing log API exposes the same information you see in the monitor, filterable by status, time window, integration flow and more. Teams commonly poll it for failed or retrying messages and push summaries into an operations dashboard or a ticketing tool. Two rules keep this healthy. First, authenticate with a dedicated technical client with only the monitoring roles it needs. Second, poll on a sensible interval with a time-window filter rather than fetching everything, so monitoring never becomes a load problem of its own. If your organisation already runs SAP Cloud ALM, its integration and exception monitoring may cover much of this without custom code.

A scenario: designing the alert for a critical interface

Quick answer: start from the business impact, define the signal and threshold that indicate it, route to a named owner with a runbook, and test the alert before go-live.

Take the order-to-cash interface that sends sales orders to a warehouse system. The business impact is clear: orders not reaching the warehouse means shipments missed. The signals are a failed message, a growing retry backlog, and, more subtly, zero messages during business hours when dozens are normal. Each becomes an alert: failures page the integration on-call immediately; a backlog above a threshold for fifteen minutes raises a high-priority ticket; a silence alert fires if no orders flow for an hour in trading time. Every alert links to the interface runbook and names the owner. Before go-live, the team deliberately breaks the connection in test to prove each alert fires and reaches the right person. That last step is the one most often skipped, and the one that matters most.

Which signals should not page anyone?

Quick answer: only page for problems that need a human now. Everything else belongs in a ticket, a dashboard or a daily report.

SignalPage someone?Better channel
Critical interface failureYesImmediate page with runbook link
Retry backlog growing on critical flowYes, after a thresholdPage or high-priority ticket
Single retry that later succeedsNoDashboard trend
Non-critical interface failureNoTicket to owning team
Certificate expiring in 30 daysNoTicket with due date
Daily volumes and processing timesNoDashboard or daily report

Getting this split right is the single biggest lever against alert fatigue. For how retry backlogs arise in the first place, see JMS vs Data Store retry, and for the full operating model, the gated monitoring guide.

What does a healthy monitoring setup look like?

  1. Message status watched per interface, with trend alerting on RETRY and FAILED.
  2. Correlation IDs on every message for fast triage.
  3. Alerts on failures, backlogs, escalations and dead-letter arrivals.
  4. Cloud ALM for landscape-wide integration and exception monitoring.
  5. Alert Notification pushing alerts to owners, with clear severity and response expectations.

Monitoring, retry and error handling are three parts of one discipline. Return to the reliability pillar, compare JMS vs Data Store retry, or go back to the platform guide.

Key takeaways

  • Watch message statuses: COMPLETED, RETRY, ESCALATED and FAILED tell you the health of every flow.
  • Alert on FAILED messages, growing RETRY backlogs, escalations and dead-letter arrivals, not just platform downtime.
  • Use the message monitor for per-message detail, logs and reprocessing.
  • SAP Cloud ALM provides integration and exception monitoring across SAP cloud and on-premise; Alert Notification pushes alerts to channels.
  • Route each alert to the team that owns the interface, with business context and a correlation ID.

Questions

What statuses should I monitor in SAP Cloud Integration?

Watch COMPLETED (success), RETRY (automatic reprocessing has started), ESCALATED (a handled error flagged for monitoring) and FAILED (processing failed with no more retries). A rising count of RETRY or FAILED on one interface is your earliest warning sign.

How do I get alerted on failed messages?

Configure alerting so FAILED messages and escalations raise notifications. SAP Alert Notification service pushes alerts to channels such as email, Slack or ticketing tools, and SAP Cloud ALM provides integration and exception monitoring across SAP cloud and on-premise components.

What is the message monitor used for?

It shows individual message processing, their status, logs and attempts, and lets authorised users inspect and, where configured, reprocess messages. It is the first place to look when an alert fires.

What is the difference between SAP Cloud ALM and SAP Alert Notification?

Cloud ALM is for monitoring actual integration flows and exceptions across the landscape, with message search, correlation and alerting on failures. Alert Notification is a delivery service that pushes alerts from BTP services and custom conditions to people and tools. You often use both: Cloud ALM to detect, Alert Notification to notify.

Is monitoring only about the platform being up?

No. Platform uptime is necessary but not sufficient. Most real incidents are specific interfaces failing while the platform is healthy, so monitor message status per interface, not just overall availability.

See what Spanovix would fix in your landscape

Bring your hardest interfaces. In a short working session we show where the agents cut failures, manual work and risk for your teams.

Or estimate your project in two minutes