Doc #52729 [Opn]: unpack() format I, L, N and V returns negative value

From: Date: Sat, 04 Sep 2010 09:40:01 +0000
Subject: Doc #52729 [Opn]: unpack() format I, L, N and V returns negative value
References: 1  Groups: php.doc.bugs 
Request: Send a blank email to doc-bugs+get-4969@lists.php.net to get a copy of this message
Edit report at http://bugs.php.net/bug.php?id=52729&edit=1 ID: 52729 Updated by: cataphract@php.net Reported by: hiroaki dot kawai at gmail dot com Summary: unpack() format I, L, N and V returns negative value Status: Open Type: Documentation Problem Package: Unknown/Other Function Operating System: Linux PHP Version: 5.3.3 Block user comment: N New Comment: OK, this is documented. The docs for unpack say: > Note that PHP internally stores integral values as signed. If you unpack a large unsigned long and it is of the same size as PHP internally stored values the result will be a negative number even though unsigned unpacking was specified. So this paragraph in pack: > Also note that PHP internally stores integer values as signed values of a machine-dependent size. If you give it an unsigned integer value too large to be stored that way it is converted to a float which often yields an undesired result. only means the internal representation of large integers may be with PHP floats and that packing those floats may yield undesired results. This isn't a valid concern for machines with 32-bit longs. If longs are 32-bit, pack doesn't make available any mode for 64-bit longs. Numbers bigger than 2^32-1 could never be handled anyway, so the only problem is that numbers between 2^31 and 2^32-1 are represented with PHP floats instead of PHP ints. However, pack arguments that are PHP floats are converted to PHP integers before packing (with a triple cast (long) (unsigned long) (long long)), and, given the usual 52-bit mantissa for doubles allows the integers in this range to be stored without precision loss, this yields the expected memory representation. If longs are 64-bit long, then it's another matter. The floats' mantissa is not long enough to store without precision loss numbers between 2^63 and 2^64-1, so the only way to pack numbers in those range to create PHP floats with an appropriate negative number. Previous Comments: ------------------------------------------------------------------------ [2010-08-30 08:12:46] hiroaki dot kawai at gmail dot com For bitwise operation, I see the problem, current implementation is so confusing. ------------------------------------------------------------------------ [2010-08-30 02:38:08] cataphract@php.net That's because when you do pack("I",4294967295) the float(4294967295) is cast into an int before the conversion. Notice (x86): $ php -r 'var_dump((int)4294967295);' int(-1) So you're actually packing int(-1) as unsigned. It seems reasonable that you receive a int(-1) back when you unpack it. Returning a float has several problems: performance, bitwise operators may not work as expected etc. Like I said, the signed/unsigned is only relevant when you may have to move the sign bit. ------------------------------------------------------------------------ [2010-08-30 02:03:56] hiroaki dot kawai at gmail dot com > unpack/pack wouldn't be inverse operations anymore. Yes it does inverse operations. Let's see the result below. -------- code <?php var_dump(bin2hex(pack("I",4294967295))); -------- result string(8) "ffffffff" pack() works fine with float(4294967295). Please note that "float" is the PHP internal representation, users should use numbers what they want. ------------------------------------------------------------------------ [2010-08-30 01:46:10] cataphract@php.net The fact that the 4294967295 integer literal results in a float is irrelevant here. What's relevant is what the documentation for pack says: > Also note that PHP internally stores integer values as signed values of a machine-dependent size. If you give it an unsigned integer value too large to be stored that way it is converted to a float which often yields an undesired result. So it does support your expected behavior. However, I'd call this a documentation bug. It doesn't make sense to unpack an integer into a float; unpack/pack wouldn't be inverse operations anymore. The current implementation still distinguishes between unsigned or signed. Since PHP uses signed longs to store integers, if the size of these longs is larger than the size of the value you're unpacking, the sign bit needs to be moved from its position in the packed value to the correct position in unpacked value. Changing this to the documented behavior would also cause probably cause a lot of breakage. ------------------------------------------------------------------------ [2010-08-29 10:22:08] hiroaki dot kawai at gmail dot com Description: ------------ According to the documentation, all I, L, N and V formats are defined to be unsigned integer. On my Linux(32bit intel), unpacking FFFFFF returns -1. As you can see the test script, we automatically use float for large integer. So the result should be float(4294967295), otherwise we get wrong number of -1. Test script: --------------- <?php var_dump(unpack("I","\xFF\xFF\xFF\xFF")); var_dump(unpack("L","\xFF\xFF\xFF\xFF")); var_dump(unpack("N","\xFF\xFF\xFF\xFF")); var_dump(unpack("V","\xFF\xFF\xFF\xFF")); var_dump(PHP_INT_MAX); var_dump(PHP_INT_SIZE); var_dump(4294967295); Expected result: ---------------- array(1) { [1]=> float(4294967295) } array(1) { [1]=> float(4294967295) } array(1) { [1]=> float(4294967295) } array(1) { [1]=> float(4294967295) } int(2147483647) int(4) float(4294967295) Actual result: -------------- array(1) { [1]=> int(-1) } array(1) { [1]=> int(-1) } array(1) { [1]=> int(-1) } array(1) { [1]=> int(-1) } int(2147483647) int(4) float(4294967295) ------------------------------------------------------------------------ -- Edit this bug report at http://bugs.php.net/bug.php?id=52729&edit=1

« previous php.doc.bugs (#4969) next »