Chrome V 124+ on MacOS - Virtual Server Access Issue

(editors note: this is not just MacOS - appears to be all Chrome browsers regardless of OS)

It appears that the latest version of Google Chrome (version 124) on MacOS (ed. any OS) has broken the above code. With debugging turned on, we get this when a MacOS client accesses a virtual server with this rule:

No SSL/TLS protocol detected ; connection is rejected (0x0000)

Can anyone else confirm this? Any idea how to fix it?  @stan_piron

----

(Editors Note: This forum post was created via a comment on the OG Article written by @Eric_Chen here:
How to use SNI Routing with BIG-IP.
That article contains ‘the above code’ and ‘the rule’ referenced here)

@Brad_Baker

I can’t speak to the issue technically but I am interested in whether you get a resolution here?
If @Eric_Chen or @Stan_PIRON_F5 can’t help right away (or at all?) perhaps a new post in the Technical Forum is in order.

I may even be able to move it automagically for you if you like

I have not heard from anyone so far. It seems this issue may be greater in scope than I realized. Its not just chrome 124 on Mac apparently its impacting Google Chrome 124, Google Chrome 125 and Google Chrome 126 on Windows, Linux, and Mac, Google Chrome 125/126 are pre-release though.

@LiefZimmerman Moving this to a new thread would be great - I would love to get some eyes on this as its impacting us greatly - more and more complaints are pouring in and we’ve had to turn off SNI on some of our more popular VIPs.

I know SNI can be implemented using F5 LTM policies but frankly we found that kludgy - its slow to work with you have to create drafts and then publish them, and set a lot of drop downs that end up being identical for all our sites.

Maybe its just our LTMs but it took 3-5 minutes to publish each policy it was SO incredibly painful to work with that we abandoned it for this irule approach which up until now has worked great. It takes second to add a new site to the datagroup. But if we can’t find a solution to this fast we may have to go back to policies :anxious_face_with_sweat:

Unfortunately this rule is quite complex - if it was simpler I might be able to edit it but I’m not sure I’ll be able to figure out what’s going on without some help.

I suspect the issue is with this change/enhancement in chrome 124:

Specifically the post above mentions:
Using X25519Kyber768 adds over a kilobyte of extra data to the TLS ClientHello message due to the addition of the Kyber-encapsulated key material.

Here’s how this is disabled in newer versions of Google Chrome:

The trouble is we can’t ask all our users to disable this flag. In order to be usable the rule needs to be modified to take into account the extra TLS ClientHello data but only for Chrome 124+

ok - I’m going to try a move this to a whole new post in the Forum…starting with the top of this thread.

(my oh my - that worked a treat)

The article above mentions “over a kilobyte of extra datain the TLS ClientHello”

I am assuming with the extra data the TCP::playload needs to be increased from 16389 to something higher.

I tried increasing the payload from 16389 to 17408 and then to 18432 and then to 19456 but that didn’t help.

So it may be more than that - maybe the pattern match for the SNI needs adjusting?

@LiefZimmerman is there any way you can edit the original post to reflect that this is not just MacOS? It’s all Chrome browsers? Maybe that would help draw interest?

I also found this: Re: Compatibility of X25519Kyber768 ClientHello

“They do not correctly handle when the ClientHello comes in in two reads. Before Kyber, this wasn’t very common because ClientHellos usually fit in a packet. But Kyber makes ClientHellos larger, so it is possible to get only a partial ClientHello in the first read, and require a second read to try again.”

Maybe this is contributing as well? I’m not sure.

I also see references to using an iRule with SSL::extensions as an alternative to the above or profiles but I can’t find any examples of how to implement that.

It looks like perhaps SSL::extensions are only valid with VIPs with a CLIENTSSL profile - which won’t work for our use case. I guess its back to looking at the irule. Can anyone help with this?  I’m not overly familiar with binary scan or TLS Client HELLO syntax.

@Joel_Moses maybe?

There is an if statement in the original iRule that checks the length of the collected payload and the indicated TLS record length contained within exactly match:

($payloadlen == $tls_recordlen+5)

…but Chrome’s record length is now much larger than that (I am seeing values between 1730 and 1850, where I used to only see smaller values around 680).

I am experimenting with removing this specific check, so $payloadlen is not compared to $tls_recordlen (or is allowed to be >=) but I don’t yet know if that will have a knock-on effect later in the script.

OK, so my $payloadlen for Chrome is now always 1380, while my $tls_recordlen is always 1750+.  I’m guessing that 1380 is the MTU, so another TCP::collect needs to be worked into the procedure to capture the entire TLS record.

Now, I remember reading somewhere that running multiple TCP::collect was A Bad Thing™, but I don’t recall the exact reasoning…

@Brad_Baker - edited.
I’m also going to escalate this through some other channels but…have you opened a ticket for it?

I’ll send you a DM so you can contact me directly (not best practice to share ticket numbers here).
* If you have the ticket number send it there
* If you don’t have a ticket - no worries; ignore that DM keep the questions/information flowing here.

Thanks

I haven’t opened a support ticket as its my impression/understanding F5 support will not troubleshoot/write/debug iRules.

@Brad_Baker   / @LiefZimmerman   / @Eric_Chen   / @Stan_PIRON_F5 / @stan_piron

I have patched my copy of the original iRule with what I think is a working fix, but I would love if someone in the community could validate it, because my knowledge of the intricacies of the TLS packet is not strong.

Essentially, I have replaced this if check in the original…

# If valid TLS 1.X CLIENT_HELLO handshake packet
if { [binary scan $payload cH4Scx3H4x32c tls_record_content_type tls_version tls_recordlen tls_handshake_action tls_handshake_version tls_handshake_sessidlen] == 6 && \
    ($tls_record_content_type == 22) && \
    ([string match {030[1-3]} $tls_version]) && \
    ($tls_handshake_action == 1) && \
    ($payloadlen == $tls_recordlen+5)} {

…with this block, which is intended to determine whether the reported TLS record length is longer than the payload we have and, if so, collect more packets until we have enough:

# Keep collecting if CLIENT_HELLO messages that span more than one packet...
if {![info exists payloadscan]} {
    set payloadscan [binary scan $payload cH4Scx3H4x32c tls_record_content_type tls_version tls_recordlen tls_handshake_action \
                                                        tls_handshake_version tls_handshake_sessidlen]
}
if {($payloadscan == 6)} {
    if {($tls_recordlen < 0 || $tls_recordlen > 16389)} {  # if we are asked to collect more than we will handle, bail...
        log local0.warn "[IP::remote_addr] : parsed TLS record length ${tls_recordlen} outside of handled length (0..16389)"
        reject
        return
    } elseif {($payloadlen < $tls_recordlen+5)} {          # if we have not collected enough yet, collect some more
        TCP::collect [expr {$tls_recordlen+5 - $payloadlen}]
        return
    }
}

# If valid TLS 1.X CLIENT_HELLO handshake packet
if {($payloadscan == 6) && \
    ($tls_record_content_type == 22) && \
    ([string match {030[1-3]} $tls_version]) && \
    ($tls_handshake_action == 1) && \
    ($payloadlen == $tls_recordlen+5)} {

This has allowed the rest of the original logic to capture the SNI server name and my services to resume operation.

However, one new thing of note, is that the resulting values in $tls_handshake_preferred_version now include an extra value.  If I log this value, I can see…

  • Firefox returns 0304 0303 0302 0301
  • Chrome/Edge with --ssl-version-max=tls1.2 return 0303 only
  • Chrome/Edge v124 return xAxA 0304 0303

The xAxA value changes every time -- so far, I have seen 1A1A, 7A7A, 8A8A, DADA and FAFA.  (I do not see a 7Fxy, as indicated in the switch statement).

I am now wondering if this is just an implementation detail of the new Chrome TLS packet, or if I/we are now reading the preferred values from the wrong position in the payload.

Is anyone able to verify the above?

The 1A1A, 2A2A, etc values appear to align with a TLS extension called GREASE: and are probably OK.

I would now just appreciate any confirmation that calling TCP::collect from within CLIENT_DATA isn’t going to cause any problems later down the line.  At most I have tried to limit to two calls (the original one in CLIENT_ACCEPTED and one extra in CLIENT_DATA).

There is one call to TCP::release at the end of CLIENT_DATA, but it is only reached if we have collected enough data across those two TCP::collect calls.

So far, in my logging, the count of bytes released does appear to match the count of bytes collected -- I’m hoping that’s all that needs to happen.

@candc - I’m no technician but I recognize work when I see it. Thank you.
(gotta say…my spam-radar went up for a second when I saw the external site link to GREASE. :sweat_smile:)

I’ll see if I can get anyone to help assess your long-term concern and the wrong payload position questions.

Cheers.

This works for me in Google Chrome 124 so tentative yay!! :slight_smile: I guess my only concern would be your question about doing to TCP Collects.

@candc I do see this:

Note: If multiple iRules attached to the same virtual execute this command, the last collect wins, meaning the earlier one will be ignored.

Here: TCP::collect

Could this be what you’re thinking of/referring to?

Yes, it was possibly that.  It is not always easy to understand the implications of those advisories.

From what I can tell, it means TCP::payload will always be updated by the result of the most-recent call to TCP::collect and, because TCP::collect can be called with different parameters (length and skip), you may lose data if two calls collide with each other.

I don’t think that is happening in my patch code above, but if another iRule was also calling TCP::collect, you could not be certain what TCP::payload would contain.

I also wondered whether there had to be a TCP::release for every call to TCP::collect, or if the one call to TCP::release would be enough.  From what I can tell, the one call is enough, but I am keeping an eye on the memory stats of my device to try to determine if I am leaking by not explicitly calling TCP::release for every TCP::collect.

(To be honest, if I had to work an extra TCP::release into the code, I can’t see where it would reliably go, and I assume I would have to therefore create my own buffer instead of using TCP::payload.  If that is the case, I’m not sure it would be anywhere near as straightforward a fix.)