Bug #76950 [Com]: trim not working

From: Date: Sat, 29 Sep 2018 23:54:35 +0000
Subject: Bug #76950 [Com]: trim not working
References: 1  Groups: php.bugs 
Request: Send a blank email to php-bugs+get-217304@lists.php.net to get a copy of this message
Edit report at https://bugs.php.net/bug.php?id=76950&edit=1

 ID:                 76950
 Comment by:         a at b dot c dot de
 Reported by:        tobias at tromm dot no-ip dot org
 Summary:            trim not working
 Status:             Not a bug
 Type:               Bug
 Package:            *General Issues
 Operating System:   Windows Server 2016
 PHP Version:        7.2.10
 Block user comment: N
 Private report:     N

 New Comment:

For what it's worth, 

preg_replace('/(^[\t\n\r\000\v\pZ]+)|([\t\n\r\000\v\pZ]+$)/u', '', $apelido);

Removes leading/trailing instances of everything that trim() removes and also everything that
Unicode considers a "whitespace" character (assumes UTF-8 encoding).


Previous Comments:
------------------------------------------------------------------------
[2018-09-29 21:40:18] tobias at tromm dot no-ip dot org

Thank you alot @requinix for the explanation.

Maybe in the future php could have a function to remove all non-breaking space if that's
possible.

Thank you again.

------------------------------------------------------------------------
[2018-09-29 21:29:15] requinix@php.net

You probably need the longer explanation.

Welcome to the world of character encodings. Nearly all of PHP's normal functions work on
bytes. You are thinking about characters. Since PHP doesn't internally manage Unicode
characters, the only reasonable way PHP can currently convert between the two is look at the byte
ranges that all ("all") character encodings can agree upon: the 0-126 range. That means
functions like trim will only deal with the bytes \n\r\t\v and space, and \0 for the fun of it, and
they will not cover anything above \x7F.

Those bytes blocking trim from reducing the entire string to "David" are above \x7F. The
exact interpretation of what characters those bytes are depends on the character encoding. In
Latin1, \xA0 (\240) is a non-breaking space, but in UTF-8 it is not a character at all - instead it
is part of a 2-4 byte sequence that represents a character (and the sequence for a non-breaking
space is \xC2\xA0 or \302\240).

There is no mb_trim function but you can use pcre_replace with \s and the /u option.

------------------------------------------------------------------------
[2018-09-29 20:59:09] requinix@php.net

Yes, it was converted. But for next time, familiarize yourself with functions like addcslashes so
that code CAN be copied and pasted.

$apelido = " \302\240 \302\240 \302\240 \302\240 \302\240 David \302\240 \302\240 \302\240
\302\240 \302\240 ";

The documentation for rtrim explicitly states what characters are trimmed by default. If you
don't like that list then provide your own.

------------------------------------------------------------------------
[2018-09-29 20:56:28] tobias at tromm dot no-ip dot org

PLEASE DONT COPY THESE CODE, IT WILL CHANGE THE RESULT COMPLETELY!!!

I just copy it and paste on my php editor and the result will change.

Use the zip file from the link I provide.

Otherwise, the space char will be converted!

------------------------------------------------------------------------
[2018-09-29 20:53:40] requinix@php.net

<?php

   $apelido = "           David           ";

   echo "Given String: \"".$apelido."\"<br><br>";
   echo "Given String with the use of php trim function:
\"".trim($apelido)."\"<br><br>";

   echo "Now we will check the ASCII table to check wich character are being
used:<br><br>";

   for ($cont=0; $cont <strlen($apelido); $cont++){
      echo "ORD: ".ord($apelido[$cont])."<br>";
   }

   $fazer = 1;

   do {
      $fazer = 0;

      //Check the beginning of the string
      if ($apelido[0] == chr(32) OR $apelido[0] == chr(194) OR $apelido[0] == chr(160)) {
         $fazer = 1;
         $apelido = ltrim($apelido, chr(32));
         $apelido = ltrim($apelido, chr(194));
         $apelido = ltrim($apelido, chr(160));
      }

      //Check the end of the string
      if ($apelido[strlen($apelido)-1] == chr(32) OR $apelido[strlen($apelido)-1] == chr(194) OR
$apelido[strlen($apelido)-1] == chr(160)) {
         $fazer = 1;
         $apelido = rtrim($apelido, chr(32));
         $apelido = rtrim($apelido, chr(194));
         $apelido = rtrim($apelido, chr(160));
      }

} while ($fazer == 1);

   echo "<br>Result String:
\"".$apelido."\"<br><br>";

?>

------------------------------------------------------------------------


The remainder of the comments for this report are too long. To view
the rest of the comments, please view the bug report online at

    https://bugs.php.net/bug.php?id=76950


--
Edit this bug report at https://bugs.php.net/bug.php?id=76950&edit=1


Thread (11 messages)

« previous php.bugs (#217304) next »