IBM Resiliency and High Availability
Tata Communications delivers Direct Link over dual interconnects to IBM by default, with gateway circuits created over both the primary and secondary NNI. The interconnects are active/active; Tata Communications applies BGP attributes so they behave as primary and secondary. IBM supports three resiliency architectures, from same-location redundancy up to multi-zone/multi-region.
Default resiliency
Every IBM Direct Link service on IZO™+ Multi Cloud Connect is delivered over Tata Communications’ dual interconnects to the IBM location, with the Direct Link gateway circuits created over both the primary and secondary NNI. The interconnects are active/active; Tata Communications applies BGP Local Preference and AS-PATH prepending so one path is primary and the other secondary, with automatic failover. A circuit is provisioned as a redundant pair, configured active/backup.
IBM’s three resiliency architectures
Tata Communications can deliver the resiliency models below on the Tata Communications network; the routing and fallback between IBM Direct Link locations on IBM’s own network should be confirmed with IBM.
Standard redundancy – same location (default)
Multiple Direct Link connections within the same physical location, typically terminated on separate devices in IBM’s network, so traffic continues if one link or device fails. This provides port- and device-level redundancy within a single facility but does not protect against a site-level failure such as a complete data centre outage.
High resiliency – different locations
Direct Link connections across geographically separated locations, delivered through independent points of presence for physical and logical path diversity. This protects against location-level outages and provider issues, and is the recommended architecture for production and mission-critical workloads.
Multi-zone / multi-region – advanced
Direct Link connections combined with workloads distributed across multiple availability zones or regions in IBM Cloud, giving independent failure domains. This provides the highest level of resiliency by integrating network redundancy with application-level failover and disaster-recovery strategies.
BFD
For sub-second failure detection on BGP sessions, BFD runs at 300 ms hello × 3, giving detection in approximately 900 ms. Apply matching BFD parameters on both sides.
Layer 2 resiliency
For Layer 2 services, two circuits (primary and secondary) are delivered by default, but failover is managed by you using routing between your CPE and IBM. See IBM L2VPN Services.
Capacity planning
When a connection failure shifts traffic to the surviving connection, that connection carries the full load. Size each connection for peak traffic, not half – if peak is 1 Gbps, provision each at 1 Gbps. Note that IBM does not support over-provisioning, so plan capacity at the connection’s actual bandwidth.
What sits with you
CPE redundancy and last-mile diversity are the customer’s responsibility: provision redundant CPE so a single device failure does not remove access, and ensure your CPEs reach the Tata Communications MPLS over diverse last-mile paths. For multi-zone/region designs, application-level failover and workload placement across IBM zones/regions sit with you.