XML_Feed_Parser and HTML security

From: Date: Wed, 16 Aug 2006 13:32:12 +0000
Subject: XML_Feed_Parser and HTML security
Groups: php.pear.dev 
Request: Send a blank email to pear-dev+get-43707@lists.php.net to get a copy of this message
There has been quite a bit of discussion in the syndication community lately about potential security vulnerabilities in feed parsers and aggregators caused by the HTML content of some feeds. Sam Ruby is one of those who has blogged about it: http://www.intertwingly.net/blog/2006/08/09/Attack-Delivery-TestSuite As I'm preparing the first stable release of XML_Feed_Parser I'd appreciate some input on how proactive that package should be in 'cleaning' HTML. At present it simply returns any HTML delivered in the feed and the expectation is that the user of the package will escape any output they get, but I'm wondering if it should be more proactive, even at the risk of a slight BC break. What I'm considering is an extra parameter to all of the methods that could potentially return infected data. By default it would process the HTML to remove potential exploits (and hence probably all javascript) but if a user passed false it would return the content as found in the feed. Alternatively that could be optional, or it could be left out entirely and kept back for a 1.1 release. Whichever way, I want there to be a clear statement about security in the docs. On a related note, I'd like to use HTML_Safe to do any parsing. Is there any word on when we can expect a stable release of that package? thanks. James.

« previous php.pear.dev (#43707) next »