Doc #73123 [Asn]: Substr documentation is wrong and misleading

From: Date: Thu, 01 Oct 2020 12:07:53 +0000
Subject: Doc #73123 [Asn]: Substr documentation is wrong and misleading
References: 1  Groups: php.doc.bugs 
Request: Send a blank email to doc-bugs+get-17937@lists.php.net to get a copy of this message
Edit report at https://bugs.php.net/bug.php?id=73123&edit=1 ID: 73123 Updated by: girgias@php.net Reported by: gregoire dot daussin at gmail dot com Summary: Substr documentation is wrong and misleading Status: Assigned Type: Documentation Problem Package: Strings related PHP Version: Irrelevant Assigned To: girgias Block user comment: N Private report: N New Comment: Can you provide an explicit example as to when substr() returns null? Because that shouldn't happen from my understanding. mb_substr() is already in the See Also section so I'm not sure what more you want for pointing in this direction, as there are various other options too, namely iconv_substr() and grapheme_substr(). The mention that a string is the same as a byte is documented on the string type page: https://www.php.net/language.types.string. Moreover, UTF-8 has *multiple* valid encodings for a single "character" (quoting because a character is a vary nebulous concept see https://utf8everywhere.org/#characters) So what is likely is that instead of having ä encoded as a single code-point (i.e. a byte for this SPECIFIC case) it is encoded as the code-point for 'a' followed by the diacritic modifier code-point '¨' thus taking at least 2 bytes. The main reason this hasn't been fixed because the proposed fix implies changing not just the documentation of substr() but every part of the manual mentioning "characters" when it talks about 1-byte encodings, which seems counterproductive, as this detail is mentioned on the string type page. Another solution would to have a note on every page about strings being byte-arrays, but that's again counterproductive. A different note could be a warning about encodings, but that's a whole different topic which is, at least in my eyes, kinda irrelevant to the topic at hand and is very complicated. And mentioning the other functions well, there is at least the mb_ variant in the See Also section as mentioned before. One possibly reasonable solution is to include an example with a multi-bytes character to highlight this. Previous Comments: ------------------------------------------------------------------------ [2020-10-01 07:13:33] balazs dot kovacs at gmail dot com Please fix this. I just ran into this issue after a very arduous debugging process which could have been cut short if there was any mention in the documentation about the mb_substr! Not to mention that my substr returned null and threw nothing while this behaviour is not in the documentation. The string I was running the function on had simple 'ä' (U+00E4) characters in them, UTF8 encoded. Also this has previously not caused any issues running through substr. ------------------------------------------------------------------------ [2018-12-30 03:44:15] girgias@php.net I will have a look at it, however, this applies to all non-mb_ functions. ------------------------------------------------------------------------ [2016-09-20 10:49:00] gregoire dot daussin at gmail dot com Description: ------------ --- From manual page: http://www.php.net/function.substr --- Substr documentation talks about characters everywhere, while the function does not cut string characters wise but bytes wise, this is especially true for UTF-8. The documentation should talk about bytes, and make a note about mb_substr. Please fix, it is confusing, I've lied to since years :(. ------------------------------------------------------------------------ -- Edit this bug report at https://bugs.php.net/bug.php?id=73123&edit=1

« previous php.doc.bugs (#17937) next »