Bug #74709 [Com]: PHP-FPM process eating 100% CPU attempting to use kill_all_lockers

From: Date: Sun, 11 Jun 2017 10:19:39 +0000
Subject: Bug #74709 [Com]: PHP-FPM process eating 100% CPU attempting to use kill_all_lockers
References: 1  Groups: php.bugs 
Request: Send a blank email to php-bugs+get-209476@lists.php.net to get a copy of this message
Edit report at https://bugs.php.net/bug.php?id=74709&edit=1

 ID:                 74709
 Comment by:         devel at jasonwoods dot me dot uk
 Reported by:        devel at jasonwoods dot me dot uk
 Summary:            PHP-FPM process eating 100% CPU attempting to use
                     kill_all_lockers
 Status:             Open
 Type:               Bug
 Package:            FPM related
 Operating System:   CentOS 6
 PHP Version:        5.6.30
 Block user comment: N
 Private report:     N

 New Comment:

Having looked into a few server, it seems the same file is used for all FPM processes. So if there
is any issue with the lock file and kill_all_lockers comes into action, there's a good chance
on most multi-pool servers of the process to start spinning.

Ideally:
- Lock file should be opened after fork in FPM, or maybe even defer extension activation until after
fork, in any case, sharing things amongst the same pool rather than the entire PHP-FPM pool set.
- A backoff on the kill so it performs once every X milliseconds instead of spinning.

I guess I have some concern on the latter though, in that it will fix the CPU but it will then have
the side effect of hiding the fact that a process is effectively deadlocked and the pool-size
reduced. So I wonder if anyone knows if the former, or similar solution, is viable?


Previous Comments:
------------------------------------------------------------------------
[2017-06-08 08:33:27] devel at jasonwoods dot me dot uk

Description:
------------
I had a spinning process eating 100% CPU on a server this morning. Attaching strace it was stuck in
an infinite loop with the following:

kill(2260, SIGKILL)                     = -1 EPERM (Operation not permitted)
fcntl(3, F_GETLK, {type=F_RDLCK, whence=SEEK_SET, start=1, len=1, pid=2260}) = 0

The process it was attempting to kill was a member of a different FPM pool, and was running under a
different user, thus the EPERM.

Stack trace of the spinning process when attaching gdb was:

#0  0x00007f33ebaf9777 in kill () from /lib64/libc.so.6
#1  0x00007f33e4b9f4a2 in kill_all_lockers () at
/usr/src/debug/php-5.6.30/ext/opcache/ZendAccelerator.c:606
#2  accel_is_inactive () at /usr/src/debug/php-5.6.30/ext/opcache/ZendAccelerator.c:661
#3  accel_activate () at /usr/src/debug/php-5.6.30/ext/opcache/ZendAccelerator.c:2192
#4  0x00000000005d80ee in zend_llist_apply (l=<value optimized out>, func=0x5d4290
<zend_extension_activator>) at /usr/src/debug/php-5.6.30/Zend/zend_llist.c:191
#5  0x00000000005d5eba in init_executor () at /usr/src/debug/php-5.6.30/Zend/zend_execute_API.c:161
#6  0x00000000005e4ae3 in zend_activate () at /usr/src/debug/php-5.6.30/Zend/zend.c:936
#7  0x0000000000582652 in php_request_startup () at /usr/src/debug/php-5.6.30/main/main.c:1647
#8  0x0000000000693f2d in main (argc=<value optimized out>, argv=<value optimized out>)
at /usr/src/debug/php-5.6.30/sapi/fpm/fpm/fpm_main.c:1927

My hypothesis based on limited reading of how this works is that it seems that a process in some
pool held an opcache lock for too long and then another process in another pool attempted to perform
an "opcache restart" of sorts, but was unable to complete it as it kept failing to kill a
process in another pool for which it has no access. There is no backoff and no "failure"
case and so it starts spinning.

Meanwhile, while this was happening, no processes in the same pool as the hung process were
responding to requests. The other pools were responding successfully. Restarting the php-fpm
completely recovered the environment.

Expected result:
----------------
No spinning process.

Should the php-fpm coordinator/parent process not handling the opcache restart? Or maybe each pool
should be using a different opcache lock so they never conflict?

Actual result:
--------------
Spinning process eating 100% CPU.


------------------------------------------------------------------------



--
Edit this bug report at https://bugs.php.net/bug.php?id=74709&edit=1


Thread (10 messages)

« previous php.bugs (#209476) next »