Re: RE: how to generate static html pages from dynamic php pages

From: Date: Thu, 09 May 2002 11:33:00 +0000
Subject: Re: RE: how to generate static html pages from dynamic php pages
References: 1  Groups: php.general 
Request: Send a blank email to php-general+get-96798@lists.php.net to get a copy of this message
> > What I'm talking about is a user goes to a nonexistent page, > > www.domain.com/products/1234 and should get a 404 error. Instead, your > > 404 page looks in the database for product 1234 and sees if it exists. > > If it does, it writes the file products/1234/index.html with the data > > from the database. Now, when the next person follows your link to > > www.domain.com/products/1234, they get your html page, instead of > > hitting PHP and the database, thus saving you some resources... > > This is cool, but I would be really hesitant to make my 404 page serve > up actual real content, even if it's only the first time the link is > hit. For the sake of search engines that might ignore links that return > a 404 error (regardless of the content at that page), I would want to > make sure my content returns a valid HTTP status code, so that search > engines know that the content is okay to index. You are missing a key concept here. You can set the return status directly from PHP, so if in your 404 handler you determine that you can serve up a real page with content, you simply send a 200 instead of a 404. There is absolutely no way for the search engine to know that it actually hit a 404 and the page was created on the fly and served up. I have built several sites that had absolutely no files, just a single 404 handler that determined what should be created. When content was added to the database the appropriate files in the filesystem were then deleted and the next request for that particular file ould hit the 404 handler which would regenerate the static file. This approach is great for applications that have a whole lot of data that doesn't change very often. One example I helped out with was a weather site. Weather reports came in every 4 hours. The entire database of hundreds of thousands of locations was thus updated every 4 hours. Now, only a tiny fraction of these locations were actually being checked by people hitting the web site, but many of these were hit *a lot*. So, previously the people who had built this would generate hundreds of thousands of static files every 4 hours even though most of these files would never actually be requested. By using the 404 trick, every 4 hours I deleted all the static weather files. The 404 handler would take a request, generate the static file for that request and serve it up as a 200. The next request for that same location would obviously not hit the 404 handler and it was served up statically. If you were to have done this using your PATH_INFO approach, you would have been hitting the database to generate the data on every single request for the same location even though you know that the data will not change for 4 hours. On really busy sites this can kill you. And in thie case the database was busy enough importing the next batch of weather reports that it didn't need these extra hits. -Rasmus

« previous php.general (#96798) next »