Observations on setting up a PEAR channel

From: Date: Sun, 04 Mar 2007 12:40:38 +0000
Subject: Observations on setting up a PEAR channel
Groups: php.pear.dev 
Request: Send a blank email to pear-dev+get-45790@lists.php.net to get a copy of this message
Having just set up my own (test, internal) PEAR channel, I thought I'd offer some thoughts and observations about the PEAR installer in general, and how it manages packages. Before I start, let me say that I appreciate very sincerely the enormous amount of work that Greg and others have put into the PEAR installer and Greg's Chiara_PEAR_Server package. It's all great. First let me state my main point of reference: RPM. For those who are not familiar with it, RPM (www.rpm.org) is a package management system for Linux operating systems and is used by various OSes including Red Hat Enterprise Linux (RHEL), CentOS, Fedora, SuSE and others. RPM, like the PEAR package.xml system, allows the packaging of metadata (about versioning, dependencies and suchlike) plus an actual piece of software (binaries, data, configuration files etc.) into an output package. RPM calls the metadata a "spec file" (the contents of which are remarkably similar to a package.xml, albeit in a different format) and the resulting package "an RPM package" (the equivalent of a PEAR .tar/.tgz output package) RPM is the base system but doesn't have any concept of "channels" ("repositories" in RPM terminology) or any features to actually support resolution of dependencies. Thus, it is commonly used alongside package management tools such as "yum" or "APT" which fulfil the missing link by enabling one to do things like "yum install [somepackage]" which will install "somepackage" and all of its dependencies, analogous to "pear install --alldeps [somepackage]". I'll mainly refer to "yum" as this is the standard tool on Fedora and CentOS, the systems I'm most familiar with. Enough about the overview, let's consider a couple of noteworthy things that make RPM/yum different to PEAR. 1. Separation of channel information and package metadata PEAR intrinsically ties a particular package file (.tgz) to a specific channel. The <channel> element has to be defined in the package.xml file and this is therefore encoded in the output package. In other words, the channel is specified at source rather than destination. In contrast, RPM spec files contain no explicit information about the repository from which they will ultimately be served. Ditto for dependencies, which are given as a package name *without* any channel information. Instead, all *channel* metadata are defined dynamically at runtime in the yum configuration file on the *end-user system* (destination), with snippets like this (simple example): [mychannelname] name = Example Repository baseurl = http://myrepo.example.com/ "mychannelname" and "Example Repository" are not special in the above configuration, and could vary arbitrarily between destination systems. Think of "mychannelname" as being like a channel alias. Separating metadata about the package itself from the transport route (channel/repository) has a number of benefits: a) It's easy to move/rename a package repository; no rebuilding of packages is necessary. b) Portability is improved; both output packages (tgz) and their respective sources can be shared/copied between channels *without changes*. There are many common uses for this, for example sharing packages between two distinct repositories - e.g. a private and public repo, where the public one might be a subset of the private one. c) It "makes sense" in a lot of ways; it is a logical separation and there is certainly an argument to be made that the mode of serving does not belong in metadata about the package d) It makes the initial packaging process simpler; a package can be built NOW and subsequently incorporated into a/some channel(s), unlike with PEAR where a channel must be set up and 'channel-discover'd before a package can even be built for that channel. This makes the barrier to entry very high for someone, especially since Chiara_PEAR_Server can be tricky to get up and running. e) It means that tricky bugs like #10254 (RPM-building specs for external channels fails) wouldn't exist. Now, just out of interest, I started to have a little look at what the technical issues might be in moving the channel metadata from originator to end-user: - The current package2 schema (and the code in PackageFile/v2/Validator.php) enforces that <channel> must be present; however, that's easy enough to change. - The REST metadata could remain exactly the same; the <c> parts would just be filled in from the actual channel that's being generated rather than from the package.xmls of the contained packages. - At least to kick off with, the dependency handling (specification of requirements) could stay pretty much the same and still include channel data. This is restrictive, but it would be a first step. (Note that removing the channel name from a dependency *does* have security implications, but not insurmountable ones - we work with it just fine in RPM land). - The main thing that would need to change (and I don't think this is huge) is that the Registry would fill in the channel identifier at install-time, *from the actual channel that a package was installed from*, rather than from the package's metadata. - We would not be able to extract channel data from static tarballs for use with things like pear make-rpm-spec, but that's OK. Are there major architectural issues I'm overlooking? Downsides? 2. Managing your own channel PEAR_Chiara_Server is a great tool. Let's get that out of the way upfront. And it has many uses. However, for simple situations (e.g.a single maintainer), it could be argued that it's overkill, not least because it requires external complexities like a database and user list. Let me explain how it works in yum land. If I have a directory full of packages (.rpm packages that is, the equivalent of .tgz output files), all I need to do is run, from the command line, "createrepo ." to create metadata for that set of packages. (That may include multiple packages and multiple versions of the same package, by the way). That creates a directory called "repodata", inside which is a number of XML files (similar to the various ones PEAR uses in REST mode) that have all the metadata about the packages in that directory. No external databases, tools or configurations are required. To set this up as a public repository (channel), *all* I need to do is start serving that directory from the web. Nothing more. Then, someone else can add ("discover") that repository on their system as described earlier and do "yum install [mypackage]". There's a beautiful simplicity and flexibility in that. Each directory is self contained, and no extra software is required to start serving a repository. Like PEAR with REST, it's also trivial to mirror, as you simply need to mirror a directory of static files. To add a new package, I just create the package, drop into the directory and do "createrepo ." again. Now, I don't think that PEAR is that much different here - the REST method of describing channel metadata is very similar to yum; I think we are just missing a command line frontend to generate the metadata rather than having to use a more complex tool like the web UI of Chiara_PEAR_server. (Again: I'm not saying that anything about Chiara_PEAR_Server is bad. It would just be nice to have a slightly simpler building-block). I think this should be possible - in theory all the XML metadata is derivable solely from the package.xml files of the packages in the directory, right? In fact if nobody else has done it, I'm tempted to code this up myself when I get the time. My use case would be something like this: pear makechannel channelname /path/to/channel for example pear makechannel pear.example.com . to create REST metadata for pear.example.com for packages in the current directory. I would structure the directory layout slightly differently, but I think that's cool as everything (including the download URL) is defined explicitly in the REST data (for the download URL, r/[pkgname]/[ver].xml -> <g>[url]</g>) So, to summarise, the PEAR Installer and RPM+yum fulfil very similar roles in their respective areas. However, as an observer, I had to jump through many more hoops to get up and running with a PEAR channel. None of them were pointless, or useless, but it seems to be me that it would be nice to lower the barrier to entry for simple cases.

« previous php.pear.dev (#45790) next »