Edit report at https://bugs.php.net/bug.php?id=71135&edit=1
ID: 71135
Updated by: mike@php.net
Reported by: iquito at gmx dot net
Summary: Random memory corruption with strings
Status: Closed
Type: Bug
Package: opcache
Operating System: Debian Jessie
PHP Version: 7.0.0
Assigned To: laruence
Block user comment: N
Private report: N
New Comment:
Some valgrind logs:
https://gist.github.com/m6w6/67f71247f5325497a59d4686ba4845ac
Previous Comments:
------------------------------------------------------------------------
[2017-09-12 12:19:15] mike@php.net
The following patch has been added/updated:
Patch Name: interned_strings_top
Revision: 1505218752
URL: https://bugs.php.net/patch-display.php?bug=71135&patch=interned_strings_top&revision=1505218752
------------------------------------------------------------------------
[2017-09-12 12:18:03] mike@php.net
Several ingredients are needed to reproduce:
* opcache.file_cache enabled
* opcache.interned_strings_buffer set low (eg. 1)
* a reasonable large codebase to fill up the interned strings buffer
* quite a bit of luck that the remaining free buffer size is low enough to trigger
I could reproducibly, but not quite reliably produce a crash or zend_mm_panic with that recipe with
eg. laravel or phan.
I've way too less knowledge of opcache to argue yet what's going on, but I'll upload
a patch which /seemed/ to fix the issue for me. Maybe someone with more insight can comment whether
I'm totally off.
------------------------------------------------------------------------
[2017-09-12 04:28:21] jjones at smugmug dot com
I failed mention it in my original posting, but we are running with "fast_shutdown=0"
already. That suggestion popped up in a debugging posting I read early on in the process.
We found a way to reproduce the SegFault's we're seeing on our PHP 7.1.9 servers today.
The servers running a normal configuration still see "php-fpm" crashes from time to time.
They are infrequent however so iterative debugging has been slow and tedious.
But we found that by reducing the "opcache.interned_strings_buffer" setting the time
between crashes was greatly reduced. Checking the resulting core dumps in "gdb" with,
set $total = (accel_shared_globals->interned_strings_end -
accel_shared_globals->interned_strings_start) + 1
set $used = accel_shared_globals->interned_strings_top -
accel_shared_globals->interned_strings_start
set $left = (accel_shared_globals->interned_strings_end -
accel_shared_globals->interned_strings_top) + 1
set $prev = accel_shared_globals->interned_strings_end -
accel_shared_globals->interned_strings_saved_top
printf "Interned Strings Total Memory: %10d\n", $total
printf "Interned Strings Used Memory: %10d\n", $used
printf "Interned Strings Free Memory: %10d\n", $left
printf "Interned Strings Previous Free: %10d\n", $prev
shows that the buffer reserved for interned strings is exhausted when each of them crashed. Going
back to the core dumps from the "normal" configurations, we found that they showed similar
stats with respect to the available interned string buffer space.
In other words, the one things the core dumps seem to have in common is that they have almost no
free space in their interned string buffers when they attempt to derefernce an invalid address
pointer. The backtraces differ, the code which triggers the SegFault, and the types of structures
with invalid pointers varied. But in the six core dumps we captured today, three from a normal
configuration and three from a system with very limited interned string buffers, when they crashed
there was very little space available in the buffer space.
- - - Core file: php-fpm719-170911:1317.dump
[New LWP 4163]
[Thread debugging using libthread_db enabled]
Using host libthread_db library "/lib/x86_64-linux-gnu/libthread_db.so.1".
Core was generated by `php-fpm: pool www '.
Program terminated with signal SIGSEGV, Segmentation fault.
#0 0x0000000000b3bc76 in zend_string_hash_val (s=0x7f97754e22b0) at
/home/jjones/ops/ops-tools/deb-build/work/php-7.1.9/Zend/zend_string.h:85
85 if (!ZSTR_H(s)) {
Interned Strings Total Memory: 4194305
Interned Strings Used Memory: 4194280
Interned Strings Free Memory: 25
Interned Strings Previous Free: 3856224
- - - Core file: php-fpm719-170911:1332.dump
[New LWP 15871]
[Thread debugging using libthread_db enabled]
Using host libthread_db library "/lib/x86_64-linux-gnu/libthread_db.so.1".
Core was generated by `php-fpm: pool www '.
Program terminated with signal SIGSEGV, Segmentation fault.
#0 0x0000000000c72bc5 in ZEND_ISSET_ISEMPTY_DIM_OBJ_SPEC_CV_CV_HANDLER () at
/home/jjones/ops/ops-tools/deb-build/work/php-7.1.9/Zend/zend_vm_execute.h:47096
47096 if (ZEND_HANDLE_NUMERIC(str, hval)) {
Interned Strings Total Memory: 4194305
Interned Strings Used Memory: 4194296
Interned Strings Free Memory: 9
Interned Strings Previous Free: 3856224
- - - Core file: php-fpm719-170911:1336.dump
[New LWP 13330]
[Thread debugging using libthread_db enabled]
xcUsing host libthread_db library "/lib/x86_64-linux-gnu/libthread_db.so.1".
Core was generated by `php-fpm: pool www '.
Program terminated with signal SIGSEGV, Segmentation fault.
#0 0x0000000000c72bc5 in ZEND_ISSET_ISEMPTY_DIM_OBJ_SPEC_CV_CV_HANDLER () at
/home/jjones/ops/ops-tools/deb-build/work/php-7.1.9/Zend/zend_vm_execute.h:47096
47096 if (ZEND_HANDLE_NUMERIC(str, hval)) {
Interned Strings Total Memory: 4194305
Interned Strings Used Memory: 4194296
Interned Strings Free Memory: 9
Interned Strings Previous Free: 3856224
------------------------------------------------------------------------
[2017-09-11 22:55:39] sroussey at gmail dot com
jjones: I wrote to nick at noodles up above about the memcache issue. Not sure why the issue is only
in 7.1 vs 7.0. We found high corruption rates in 5.6.x where x>9 as well. Never figured them all
out, but fast_shutdown=0 helped a bit. The 7.0.x series has been the most stable for us.
That said, I really don't recommend rsync as a deployment strategy. Feel free to message me for
other ideas for a high load site.
------------------------------------------------------------------------
[2017-09-11 10:24:23] jjones at smugmug dot com
Re: coming up with a reproducible use case that triggers a crash. We're working on that, since
it would be the ideal way to find/fix the problem. But no luck so far. So that's a "work
in progress".
Re: Valgrind, I thought about that and wasn't sure how usable the system would be in terms of
performance with that extra overhead added to the "php-fpm" processes. But at this point
I think it's worth a shot.
Re: disparity in the crash dump outputs when the root cause it memory corruption. Yup! I have two
more core dumps so far and I haven't found anything that links them together yet aside from
impossible values for various address pointers. I'm going to dig into them more and keep
searching though.
------------------------------------------------------------------------
The remainder of the comments for this report are too long. To view
the rest of the comments, please view the bug report online at
https://bugs.php.net/bug.php?id=71135
--
Edit this bug report at https://bugs.php.net/bug.php?id=71135&edit=1