Bug #67796 [Opn->Dup]: php-fpm: repeatedly spinning many many workers
Edit report at https://bugs.php.net/bug.php?id=67796&edit=1
ID: 67796
Updated by: nikic@php.net
Reported by: kenny at kennynet dot co dot uk
Summary: php-fpm: repeatedly spinning many many workers
-Status: Open
+Status: Duplicate
Type: Bug
Package: FPM related
Operating System: Debian
PHP Version: 5.5.15
Block user comment: N
Private report: N
New Comment:
Another manifestation of bug #73342, marking as duplicate.
Previous Comments:
------------------------------------------------------------------------
[2014-09-07 06:57:04] kenny at kennynet dot co dot uk
This is the same issue as bug: #61558
My use-case was ssh using a ControlMaster file (as described in that bug report).
In this scenario ssh sets stdin to O_NONBLOCK but never removes it.
Without a ControlMaster, ssh does reset (remove) the O_NONBLOCK flag on stdin.
------------------------------------------------------------------------
[2014-08-13 05:02:19] 64438136 at qq dot com
Excuse me, how to solve this problem?
------------------------------------------------------------------------
[2014-08-06 13:39:11] kenny at kennynet dot co dot uk
Update:-
I can 100% reproduce the problem and I'm now not sure if this is a PHP, PHP-FPM or not.
Within php if you exec(), the process you exec inherits stdin from the -fpm worker (e.g. the
listen_socket).
If that process then sets O_NONBLOCK this causes the -fpm issue described.
You can reproduce with this C file (nonblock.c): http://pastie.org/9450459
Exec'd from a php script called from a fpm worker:-
"""
<?php
header("Content-type: text/plain");
passthru("/path/to/nonblock");
"""
Just refresh the page to cause (and fix) the issue at will.
------------------------------------------------------------------------
[2014-08-06 11:28:34] kenny at kennynet dot co dot uk
Description:
------------
After php5-fpm has been running for an amount of time we start to see log messages spinning round
rapidly many many times of second of the form:-
"""
[06-Aug-2014 12:18:46] NOTICE: [pool pool-1] child 10141 started
[06-Aug-2014 12:18:46] NOTICE: [pool pool-1] child 9764 exited with code 0 after 0.567466 seconds
from start
"""
(A side effect of this is that the scoreboard is updated with "idle++" but if no work is
processed that idle counter is never decremented so it continues to increase)
Attaching gdb and following the fork to the child one can observe that the accept() call in
fcgi_accept_request() returns -1 with errno=EAGAIN.
A result expected from non-blocking sockets but not blocking sockets, further:-
gdb> print fcntl(listen_socket, 3)
... 2050
Which is O_RDWR | O_NONBLOCK.
So it appears that *somehow* the listen_socket is being put into non-blocking mode and hence:-
* The loop around fcgi_accept_request() aborts immediately with exit_status=0.
* The child subsequently exits.
* fpm_children_bury() processes this as restart_child=1
* Round and round it goes spinning up new new children which exit immediately if there is nothing to
accept().
As a test / proof of the above this is fixed by:-
gdb> print fcntl(listen_socket, 4, fcntl(listen_socket, 3) &~ 04000)
... 0
Well, fixed until the next time the socket somehow gets put into non-blocking mode, I've not
yet managed to isolate exactly how that is happening.
Test script:
---------------
N/A
Expected result:
----------------
N/A
Actual result:
--------------
N/A
------------------------------------------------------------------------
--
Edit this bug report at https://bugs.php.net/bug.php?id=67796&edit=1
Thread (6 messages)