Comments on non-unique naming convention for closures
| From: | Terry Ellison | Date: | Wed, 20 Nov 2013 16:41:46 +0000 |
| Subject: | Comments on non-unique naming convention for closures | ||
| Groups: | php.internals | ||
| Request: | Send a blank email to internals+get-70237@lists.php.net to get a copy of this message | ||
The following bugs relate to this discussion:
#64291 Indeterminate GC of evaled lambda function resources
#65915 Inconsistent results with require return value
I don't want to discuss these or the specific fixes here, since I can work up a fix and discuss them in the bugrep. However,
my one-sentence Q is "should we replace the naming convention for closures with a truly unique one?" What I would like is some feedback / guidance / discussion on the general architectural issue which underlies the reasons for these bugs occuring in the first place.
* The associated closure functions are compiled during the INCLUDE_OR_EVAL opcode that included / eval'ed the originating PHP source, and are named according to the convention "\0{closure}$name$addr" where the name is the resolved filename of the source where the function was defined and the addr is the absolute address of the function definition within the memory resident copy of the source during the compilation process. The compilation creates an entry in the GC(function_table) for this function
* Closure objects are instantiated during the execution ZEND_DECLARE_LAMBDA_FUNCTION opcode at runtime. This handler invokes the built-in CTOR for the closure bind the corresponding function,
* Closure objects are destroyed by the closure DTOR in line with the usual PHP scoping rules.
* There is no specific deletion function for the GC(function_table) entries other than normal request rundown, because the compiler and RTS can't safely determine if a closure function will no longer be required for further closure, e.g.
for ($array as $i) {
$f = function($x) { ... };
...
}
* The "\0{closure}$name$addr" scheme does provide a sort of dirty reuse in that if another function is compiled with the same filename and at the same source address (which can happen for logically different variables due to storage reuse), then the current implementation simply overwrites the earlier function definition. Though the rationale for this isn't documented, I assume that this is because it is _usually_ safe to assume that the earlier copy is no longer required. However, it is possible to construct cases where this assumption is false -- especially as is the case of OPcached functions where the name can persist for the life of the SMA rather than just one request.
What I am suggesting is that this convention should be changed for something like "\0{closure}$name$ctime$index" or even just "\0{closure}$uuid" using the standard UUID algo. Reactions? Comments?
The only side effect that I can see is that in circumstances where the overwrite *does* replace stale functions then the function DTOR will no longer be taking place leading a growth in the function table with the associated memory overhead.
Thanks for any feedback
Terry