Bug #77937 [ReO]: preg_match failed

From: Date: Thu, 25 Apr 2019 17:03:31 +0000
Subject: Bug #77937 [ReO]: preg_match failed
References: 1  Groups: php.bugs 
Request: Send a blank email to php-bugs+get-220615@lists.php.net to get a copy of this message
Edit report at https://bugs.php.net/bug.php?id=77937&edit=1

 ID:                 77937
 Updated by:         cmb@php.net
 Reported by:        v-altruo at microsoft dot com
 Summary:            preg_match failed
 Status:             Re-Opened
 Type:               Bug
-Package:            PCRE related
+Package:            *General Issues
 Operating System:   Windows 10
 PHP Version:        7.3.5RC1
 Assigned To:        cmb
 Block user comment: N
 Private report:     N

 New Comment:

This issue is neither directly PCRE nor testing related.  Consider
the following script:

    <?php
    var_dump(17.4);
    var_dump(setlocale(LC_ALL, 'pt_PT'));
    var_dump(17.5);
    var_dump(ctype_alpha(224));
    ?>

This outputs on my Windows system:

    float(17.4)
    string(5) "pt_PT"
    float(17,5)
    bool(false)

The first three lines indicate that pt_PT is properly supported,
but the failing ctype_alpha() shows that it is not really.

The following C program confirms that the issue is not directly
related to PHP:

    #include <stdio.h>
    #include <ctype.h>
    #include <locale.h>

    int main()
    {
        struct lconv *lc1 = localeconv();
        char *loc = setlocale(LC_ALL, "pt_PT");
        struct lconv *lc2 = localeconv();
        int alpha = isalpha(224);
        printf("%s %s %s %d\n", lc1->decimal_point, loc, lc2->decimal_point, alpha);
        return 0;
    }

Outputs on my Windows system (when built with VC15):

    . pt_PT , 0

Again, ctype fails to properly recognize the locale (which is the
reason for the failing test, since PCRE2 calls ctype functions to
build the character tables).

If I build with VC11, I get:

    . (null) . 0

Apparently, AppVeyor behaves either like this, or it properly
recognizes pt_PT for the ctype functions.


Previous Comments:
------------------------------------------------------------------------
[2019-04-25 10:03:38] requinix@php.net

Hmm, yes, it seems Windows will quite happily accept any "language" or
"language_country" string regardless of whether either part exists, as long as the
language code is 2 or 3 characters.

var_dump(setlocale(LC_ALL, "xjq_ASDF")); // returns xjq_ASDF
var_dump(setlocale(LC_ALL, "0")); // still xjq_ASDF

FFS.

So for maximum portability it seems you have to list Windows-specific strings before the normal
strings. Or at least the codes it accepts before any short ones.

setlocale(LC_ALL,
  "Portuguese_Portugal.28591", // windows okay (28591 is the codepage for ISO 8859-1),
linux ignored
  "Portuguese_Portugal",       // windows okay, linux ignored
  "Portuguese",                // windows okay, linux ignored
  "pt_PT.ISO8859-1",           // windows ignored (bad codepage), linux okay
  "pt_PT",                     // windows okay (wrong), linux okay
  "pt"                         // windows okay (wrong), linux okay
);

------------------------------------------------------------------------
[2019-04-25 09:14:28] cmb@php.net

Thanks for reporting!  I can reproduce the *test* *failure*. The
problem is that setlocale()[1] claims to support "pt_PT", but
actually it does not.  Actually supported locales would be "pt-PT"
and "portuguese".

I'm not sure yet what to do about this.  Simply fixing the test
case for Windows would be an option, but that would not fix the
underlying issue which may affect existing userland code.

[1] <https://docs.microsoft.com/en-us/cpp/c-runtime-library/reference/setlocale-wsetlocale?view=vs-2019>

------------------------------------------------------------------------
[2019-04-24 22:24:29] a at b dot c dot de

Incidentally, the test cited in the original report uses the string
"aàáçéè", not Hebrew characters.

------------------------------------------------------------------------
[2019-04-24 22:20:37] a at b dot c dot de

Adding the /u modifier to the pattern would help (assuming the source is encoded in UTF8 - you
can't even SAY "aאבחיט" in ISO8859-1).

------------------------------------------------------------------------
[2019-04-24 20:45:09] requinix@php.net

Last I knew Portuguese does not cover Hebrew characters.

------------------------------------------------------------------------


The remainder of the comments for this report are too long. To view
the rest of the comments, please view the bug report online at

    https://bugs.php.net/bug.php?id=77937


--
Edit this bug report at https://bugs.php.net/bug.php?id=77937&edit=1


Thread (9 messages)

« previous php.bugs (#220615) next »