Re: [RFC] Throw error for passwords longer than 72 bytes in password_hash() with bcrypt
| From: | Andrey Andreev | Date: | Wed, 30 Sep 2026 20:01:05 +0000 |
| Subject: | Re: [RFC] Throw error for passwords longer than 72 bytes in password_hash() with bcrypt | ||
| References: | 1 2 3 4 5 6 7 8 | Groups: | php.internals |
| Request: | Send a blank email to internals+get-132731@lists.php.net to get a copy of this message | ||
Hi Tim,
On Wed, Sep 30, 2026 at 9:03 PM Tim Düsterhus <tim@bastelstu.be> wrote:
> >
>
> https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.800-63B-4.pdf#subsubsection.3.1.1
>
> You are correct in that the NIST standard specifies that:
>
> > 9. Verifiers SHALL request the password to be provided in full (not a
> > subset of it) and SHALL verify the entire submitted password (e.g., not
> > truncate it)
>
> However: The NIST standard applies to the “Verifier” (i.e.. the
> application), not to the primitive. While
password_hash() is
> certainly
> intended to be usable as-is, an application requiring strict compliance
> with all SHALLs and SHOULDs of the NIST standard already needs to
> perform some steps that cannot be solved by password_hash(),
> such as
> verifying that known-insecure passwords cannot be used. And when trying
> for pedantic compliance with NIST you cannot use BCrypt at all, since
> BCrypt technically is not in the list of NIST-approved hashing
> functions. For that you need an approved hash function, and my latest
> knowledge is that neither BCrypt nor Argon2’s BLAKE2b primitive are
> acceptable. In practice this means using hash_pbkdf2().
I pointed to NIST SP 800-63 as the root source, but not the only
standard. If I pointed to the OWASP ASVS for example, most of what you
wrote would be irrelevant, but my basis - the correct implication of the
bytes vs characters distinction - is still proven. All of the math
arguments you've brought up previously are eroded by that, and I really
must say it - deceiving developers and end-users about behavior is not a
math problem to begin with.
What you're technically correct on is the Verifier vs "primitive" (not
really a primitive in the true sense of that term). However, obfuscating
bcrypt's deficiencies is actively reducing developers' capacity to
compensate for that, and leaving them with a false sense of security.
If we ignore all that, implementing the RFC would mean that BCrypt
> hashing with PHP would start to deviate from a combination of two
> different requirements:
>
> > 2. Verifiers and CSPs SHOULD permit a maximum password length of at
> > least 64 characters.
>
> and
>
> > 4. Verifiers and CSPs SHOULD accept Unicode [ISO/ISC 10646] characters
> > in passwords. Each Unicode code point SHALL be counted as a single
> > character when evaluating password length.
>
> The RFC would thus trade a SHALL violation for a SHOULD deviation that
> affects exactly the users using non-ASCII character sets.
>
I don't see an added SHALL violation; the RFC is only making an
already-present limitation visible.
And while the effective length is reduced for e.g. CJK characters, they
> are also much more information dense than ASCII characters. An English
> word taking up 5-7 bytes can often be represented as a single 3-byte CJK
> character. Of course the RFC also does nothing about this: Instead of
> reducing the effective length, inputs exceeding “the effective length”
> are rejected, forcing the user to pick a shorter password. The
> achievable security ceiling is the same - and high enough.
>
Might be true for CJK; not true for Cyrilic, Greek - halved in length and
strength no matter how you look at it. And I'm not mentioning others simply
because I'm not familiar with them, but there are multiple.
> The NUL byte point is moot, because NUL bytes are already rejected
> during hashing. For that I agree with the ValueError / check, since the
> truncation on NUL is much more catastrophic - and because actually
> including a NUL in a legit user-provided password is much more
> complicated than exceeding a 72 byte limit. In practice the Internet has
> also converged on “just using UTF-8”.
>
I admit I keep forgetting that the NUL byte thing was already patched. The
point isn't entirely moot still, but not worth arguing third-order effects
and we have plenty of other disagreements, so be it.
> I don't necessarily disagree with the BCrypt truncation being a problem,
> but in this case the cure is worse than the disease.
>
As long as you recognize the problem, I don't think we'd be far apart. I
absolutely recognize the BC break as a major concern that might take higher
precedence. Until now though, your position has been a lot closer to "no
problem here at all" than "I draw the line at breaking existing
applications".
Cheers,
Andrey.