Re: File_Repository Class Proposal

From: Date: Sat, 21 Sep 2002 01:33:29 +0000
Subject: Re: File_Repository Class Proposal
References: 1 2 3 4 5  Groups: php.pear.dev 
Request: Send a blank email to pear-dev+get-9271@lists.php.net to get a copy of this message
Mike McCallister wrote:
There is an interesting idea. I have not thought of that yet - probably because I have not required that type of functionality. Sounds good to me - I guess I should add that?
Or do I pearify it first, add it to CVS and then you add it - how does that work around here?
yeap that is a good way to do it - pearify + add to cvs, then if people say 'can I add this to it'- just say yah/nah etc.
Wouldn't there be issues with Windows and the default location of "mimetypesFile"? Not that I care to much for Windows, but since it will be in PEAR isn't it supposed to run on all OSes?
I did consider including a real mime.types file as part of the package - then use file (dirname(__FILE__)./mime.types")
I don't see in HTML_Template_Flexy where mimetypesFile is specified in the options property.
its not from that, it's from another package that is still under development - A mostly simple URL=>class mapper, called HTML_FlexyFramework.
I am assuming that PEAR::getStaticProperty('HTML_FlexyFramework','options') loads FlexyTemplate and mimetypesFile gets set according to some config?
yeah the point of that line is so you can configure the location of the mime.types file by putting it in an ini file, then loading like the first example here. http://pear.php.net/manual/en/packages.database.db-dataobject.configuration.php PEAR::getStaticProperty('HTML_FlexyFramework','options') would relate to the array from the [HTML_FlexyFramework] section of the ini file.
I primarily use the File_Repository for session files on only three sites where they get over 1000 new sessions per minute (the only time I use it is when performance is an issue) - so compile time is a consideration if it is going to be loading. Therefore, adding serving IMHO should probably be a subclass file of File_Repository so that the serving stuff is only loaded if needed. What are your thoughts?
Sounds fair - File_Repository_HTML ... for lack of any better ideas...
Also, is this a E_ALL workaround? !@$options["mimetypesFile"] - just curious about why there is a "@" in there.
yeap, its a hide errors trick - eg. removes the need for (isset($options["mimetypesFile"]) && !$options["mimetypesFile"])
Mike Alan Knowles wrote:
have you thought of adding a 'serve' method to it?, which basically outputs the file + the correct mime header function serve($mimetype=NULL){ if ($mimetype===NULL) {
       $mimetype = $this->getMimetype($this->filename);
....
    header('......
    fopen(...)
    while (!foef()) {
       $f =fgets(.. 4096);
       echo $f;
    }
fclose(...) } this is a little hack to get around the lack of mimetype extension in current versions of php, function getMimetype($filename) {
      $options = PEAR::getStaticProperty('HTML_FlexyFramework','options');
      if (!@$options["mimetypesFile"]) {
          $options["mimetypesFile"] = "/etc/mime.types";
      }
      $bits = file($options["mimetypesFile"]);
      $map = array();
      foreach($bits as $line) {
          $line = trim($line);
          if (!$line) {
              continue;
          }
          if ($line{0} == "#") {
              continue;
          }
          $parts = preg_split ("/[\s]+/", $line);
          if (count($parts) < 2) {
              continue;
          }
          for ($i=1;$i<count($parts);$i++) {
              $map[$parts[$i]]  = $parts[0];
          }
      }
      if (preg_match('/.*\.([a-z]+)$/', strtolower($filename), $args)) {
          if (@$map[$args[1]]) {
              return $map[$args[1]];
          }
      }
      return "application/octet-stream";
} Mike McCallister wrote:
My bad - forgot to post link to the code! Please keep in mind that this class is part of a larger framework (of about 50,000 lines of PHP code) - so there will be references to stuff that is not a part of this class. I will obviously remove that stuff before posting (if posting is deemed a good idea) to CVS. Anyways, here is the class: http://caffeine.contactdesigns.com/~jolt/Repository.php Please keep in mind that the file locking uses flock() and therefore will not work on network file shares. Let me know if I should take the time to convert it over to PEAR - I prefer the LGPL if that has any bearing on anything. Heh - there is a bug on line 271 that will prevent it from running on Windows (which isn't a problem around here ;) I will fix that too. Mike Alan Knowles wrote:
Excellent :) - I've been looking at doing the blobs stuff for my midgard db store emulation layer - (this is almost exactly the same method midgard uses, for blobs), plus you solved the issue of storing filename, when you dont have a database, or redirect to the file on a cgi server. +1 here, although posting the code might be a good idea :) Regards Alan Mike McCallister wrote:
Greetings, I constantly feel guilty for using PEAR code all the time without having contributed more than bugfixes. So here is a File_Repository class you guys can have if anyone is interested (if not, you can't blame me for not trying ;). I'd have to Pearify it of course and add PEAR error handling and clean it up a little (although it is reasonably clean). Currently, it does not autogrow the repository - it has to initialize it first. Anyways, let me know if there is any interest. Here are the docs from the top of the class file (should give you an idea what it does): * This class is designed to deal with a large store of files that need to be * accessed quickly. The majority of modern filesystems do not deal well with * accessing files in a directory where there are many files (say 5000+). Why? * Well basically when the OS wants to a get a file handle, it will ask the * directory which inodes that file lives on. When a directory
has inode
* information on MANY files, it can take longer (sometimes much longer) to find * out inode information for a single file. In reality there is a bit more to it * than that but that is the basic idea. Some filesystems (i.e.
ReiserFS
* http://www.reiserfs.org/) don't have this problem as they use more advanced * algorithms for looking up inode data (usually some form of binary tree). So * if you are running one of these cool filesystems this class will do you * NO GOOD WHATSOEVER. This class is used to make file access
quick on a
* filesystem regardless of what kind of filesystem it is. How does it do this? * Quite simply it creates a directory hierarchy (usually with a depth of 2). * Each level has 62 directories in it (A-Za-z0-9). Therefore a 2
level
* hierarchy has 622 or 3844 directories. Since file access is still * reasonably fast on directories with less than say 500 files, a two level * directory hierachy (aka repository) can store 500*3844 or
1,922,000
files and * still have speedy file access. I don't recommend a three level repository * EVER (623 or 250,047 directories) - if you have this many files, you NEED * to change your filesystem. Of course, you can still use this class on the * nifty filesystems as a means to organize the files - this can be a good thing * since it can be very annoying to "ls" in a directory only to have 2 million * files returned ;) * * So what does this class do for you? Well, first it can create the repository * for you which is good because creating 3844 directories by hand could really * suck. Next, it gives you three important methods for accessing/updating files * in the repository: open(), store(), and retrieve(). These methods abstract * away the fact that there is a directory hierarchy. In other
words,
by using * them, it is as is you were manipulating files in a single
directory.
* * This class was written specifically for two primary uses (although others * exist whereever you have the need to store A LOT of files): a more robust * session file storage and retrieval system for sites with medium to high * traffic AND storing files associated with database records. WHY store files * associated with database records when you can just store them as a BLOB type * with that record? First, the filesystem is a much more convenient way to * store files (i.e. don't need SQL to access file). Second, on servers that are * running many databases that get accessed frequently (i.e. our web servers), * returning BLOBS can wipe out your query/index memory cache
therefore
* negatively affecting other databases on the same server. * * It is important to understand how we store/organize files in the repository. * It has one key strength and one key weakness (with a work
around). By
* default, it will will store a file based on each beginning character up to the * max levels (2). So a file named "test.txt" would be stored in /root/t/e. * This is a good thing because it is easy for developers to track down which * directories contain files since it makes intuitive sense. This
way of
* organizing files also has a weakness - storage location is ONLY based on the * first N characters. So if each file always started with "prefix" all files * would end up in /root/p/e and therefore there would be no
advantage
to using * this class. There are two ways around this problem. The first is to set * auto_md5 to TRUE. This will prefix each filename with a 32 character MD5 * digest for example: dd18bf3a8e0a2a3e53e2661c7fb53534_test.txt. The digest * is sufficiently random that you will get an even spread over the repository * and each digest is unique to the filename and will always be the same for the * same filename. While this will give you the best spread, it is not convenient * to look files up since you have to compute the digest to figure out what the * first two characters are. The other way is to set prefix_seed to something * other than 0. In the case of all files starting with "prefix" a file called * "prefix_test.txt" would be stored as if it were test.txt by
setting
* prefix_seed to 7. Of course, this is really only useful when all files have * a common prefix. The MD5 method is the best way to store the files as it is * not affected by common names, prefixes or patterns in filenames. * * When you specify root during object initialization is MUST NOT HAVE A TRAILING * dir_delim and it MUST BE ABSOLUTE. If it does, this class will not function * correctly. * * Usage: * * $rep = new CDS_File_Repository(array('root' => '/path/to/dir')); * $data = $rep->retrieve(array('filename' => 'test.txt'));


« previous php.pear.dev (#9271) next »