F5 health monitors intermittently failing.

Hi All,

I am facing a wired issue with some (say 2 F5) different environment. Issue is pool is going intermittently flapping due to health monitoring (icmp, tcp & udp) failing  with error “No successful responses received before deadline.” Both F5 environment are running on version - 17.5.1.3

One interesting thing observed was that one F5 environment this health monitoring failing was observed on standby F5, while there were no issues observed on active F5. Please let me know what can be the actual cause of this issue & how to resolve this.

Hello @Preet_pk

F5s perform health checks form their self IPs.
So working from one while not from the other might means that not both IPs are allowed for FW for example.
Have you checked it?

BR,
Kostas

Why three monitors? That is 3^2 , 9, possible states. You have to tell the F5 how you want to respond each of the 9 states.

I normally use only one monitor. There are only two states up and down.

Then I use the monitor that is highest up the OSI model. Example HTTP or HTTPS for a web server.

More then one monitor can be used when something other member needs to be monitor. The default behavior for a monitor is to use the pool member IP address and service port. You use a custom monitor to check the health of a different service port on the pool member or completely different IP and port.

Have you asked the administrator for the pool member if the health monitors are getting blocked by something on the pool member?

Is there a firewall or network filtering between the LTM and pool members?

Is there any network static routes on the LTM that is causing the health monitors to send with an unexpected SelfIP address as the source IP address?

Hi @mwolf ,

Just to clarify: this issue is not limited to a single pool — it’s affecting multiple pools, each configured with health monitors across ICMP, TCP, and UDP (not a case of 3 monitors on one single pool).

Also, as I mentioned earlier, this is occurring only on the standby F5, not on the active F5.

Could you please check whether this is related to the bug described here, and if so, whether there’s a temporary workaround available apart from upgrading:

You could use three different monitors to figure out 3 different facts.

  • ICMP doesn’t work - could be routing, fw or localhost fw
  • TCP doesn’t work - icmp works, could be fw or localhost fw or service
  • HTTP doesn’t work - icmp and tcp works, service is up, most likely an issue with the app itself

This will hint you, to whom you have to talk first.

F5 customer support should be able to help determine whether you are being affected by a bug.

Multiple monitors add complexity that is not required in most situations.  The are exceptions any rule.

It’s very rare that the ICMP monitor is needed in addition to a protocol monitor. On most BIG-IP it’s easy to access the Unix shell for troubleshooting.  You can manually test the IP forwarding path to the pool member via ICMP echo requests via the ping program or ip routing and mtu tracepath.

ICMP is not good for monitoring pool members that are not on an IP network that is not directly connected to the BIG-IP.  ICMP data-grams can be dropped when there is any network congestion. I try to limit the ICMP monitor to only directly connected pool members. Example the members of a gateway fail-safe pool.

“Simplicity is the ultimate sophistication.” – Leonardo da Vinci.

“Keep it simple, stupid.” - Kelly Johnson

“Perfect is the Enemy of Good.” - Voltaire