Re: Function caching...

From: Date: Sun, 10 Dec 2000 01:18:09 +0000
Subject: Re: Function caching...
References: 1 2 3 4  Groups: php.dev 
Request: Send a blank email to php-dev+get-40715@lists.php.net to get a copy of this message
Kristian Koehntopp wrote: > Ron Chmara wrote: > > > The current performance limitations in PHP come from multiple > > > design limitations. > > Keep in mind that some limitations are apache, some are the > > web, some are PHP. > These limitations ARE. I do not care where they come from. Well, you should. Complaining to the PHP lists that the web is stateless is a bit silly. PHP code can help in some ways, but it cannot overcome all of the problems you have. > There > exists solutions around these limitations. Tried solutions. They > usually work the way I described them in my previous mail to > Zeev. They do work that way for a reason. I explained why, too. Right. It sounds like you have tried them, discovered that they have some serious issues involved in them, and are now hoping somebody else has better solutions. Many of the problems, however, are issues where you have already discovered that the solutions are slow, inefficient, or require a completely different language than PHP for scripting. > > If you want speed, abstractions are the wrong approach to take. > > Scripting languages, creating abstractions, to serve over web > > pages?... ouch. Write it in C, not as objects in PHP. > That buys me what? I am still running under the wrong UID. Web ports know no user. You can hack around these, with *major security flaws*, because of the nature of web connections.... PHP can internally add duplicate functions, apache can add suexec features, but hiding this problem under a different rock does not solve the problem. A web server is not a desktop, you have no guarantee that it's the same person at a keyboard and mouse. If you wish to allow web users greater access, you can do that. With any development environment, this can be as dangerous as you want it to be. > I > still lose my state variables at the end of the page, including > all file and network (database) handles. The web is stateless. PHP does allow re-use of database connections, but not file handles. Of course, maintaining open file handles is dangerous (think 16000 hits... instant OS handle overload). > I am still unable to > have persistent ressources, and I am still unable to share > ressources among control flows as long as these flows of control > are in different processes as they are in the Apache model (The > current Apache model is not broken, it is very resistant against > failure. The brokenness comes from the attempt to run large > amounts of code inside the Apache process model instead of > creating dedicated persistent code execution processes called > application servers). Well, PHP and Apache were written for the web. The web is stateless. The models added on top of this stateless environment to emulate state can be more, or less, efficient, depending on how they're written. PHP was not written with the kind of state emulation you desire, which is *why* it's easier to use than something else with lots of code overhead to manage state. This is why adding libraries to PHP to manage state will always slow it down, because of the additional work invovled in imposing an artificial state into a stateless environment. > > > this network of objects is rebuilt from scratch in each > > > instantation of your script, it will take significant amounts of > > > time. > > I can't think of a good reason to design a fast application in > > this way. > You have been doing small projects without much backoffice > integration The M$ product? Of course not. Trying to build a working app on top of a inherently broken one is a real problem. If you mean a complex, 400+ sites, multi-level, 5000 active concurrent users, 3 db+ldap, 3 physical sites, 19 machines, environment, that's my current main project. > and much code reuse in the past, and these were not > sensitive to security issues. Hah! Nope. It's all in the implementation. I've replaced 400 lines of security code with 8 (going from some php-to-db auth to LDAP), I've done db-redesign to implement a db model in a true RDBMS (MySQL was slowing them down, 300% improvement) etc. Good code is unrelated to the amount of code lines used, or the abstraction layers. > I can think of many reasons to > design applications that way, and so can Ulf, There are also good reasons to avoid it, especially when using PHP, or when needing a fast application, or when programming for the web. > and so can the > Twisd people (they worked around their setup time penalty by > going to C++, but still they lose context at the end of each > page). The web is stateless, so this is appropriate. > Going to C or C++ is not a solution, if there is a chance to > have a PHP framework with a different execution model that can > do the same. PHP code is much, much cheaper per line than C++ > code. The development cost is always less expensive for something with less features, with less complexity. > > Abstractions speeds up development, maintenance, and ease > > of use, at the _cost of execution speed_. > You are using abstactions wrongly, or in the wrong execution > context (for example, you are throwing initalized execution > contexts away at the end of each page when they are perfectly > reuseable). No, it's a question of layers. Adding layers mean that there is more code to pass variables through. Example: Variable passed to page. Variable passed to DB abstraction prep layer. Variable passed to DB abstraction execution layer. Variable passed to db script layer (PHP). Variable passed to db driver. Variable passed to db. Compare to: Variable passed to page. Variable passed to db script layer (PHP). Variable passed to db driver. Variable passed to db. Regardless of execution context or resuse, you are adding uneeeded layers of interaction that a CPU must handle, that need memory space, the require processing time. > > > Even if you just revive the dead object network from > > > serialization this will slow you down. > > Right. So don't use the objects. > And that does buy me what? faster code. > Setup times from the serialization of > the equivalent array network serialization and unserialization > AND it adds namespace management problems. This is hardly a good > design decision. Speed is not always easy, or pretty. If you are designing for ease of programming, that is a different goal than speed. For maximum speed, you have have to reduce your code to the bare minimums, you have to make different design decisions. > > OO is simply slower in > > execution because each object is loaded down with dead-weight, > Done properly, an object is exact the same size as a hash, plus > a single pointer to its class description. This is hardly > dead-weight. Also, you need a class description listing all > possible slots, and all object functions with their signatures > and function pointers. You need such a class description exactly > once, as it is shared between instances of that object. It's not the size of a single, pristine, object that is the problem. It's the overhead of object management, the overhead of creating and maintaining the separate memory spaces, the workflow time in creating the objects, the problems of code transparency which reduces re-use (which encourages object multiplication), the additional overhead added to OO for rarely used methods, etc. etc. Ideally, objects are almost identical. If this was the case, you would have *no problem* in getting rid of them. :-) It's the additional features and capabilities of OO design that changes the execution speed, not how many bits of memory a hash takes up. > > Quite simply, designing applications for the web is completely > > *unrelated* to design principles for desktop applications. > Not at all. It is in fact exactly the same once you get the hang > of it. Oh, so you _aren't_ having problems with state, or with compile-on-each-run instead of compile-once-then-distribute? ;-) Your complaints are almost all related to trying to use desktop application design techniques on the web. > That's what PHPLIB has tried to teach you for the past > few years, and that's what you can learn von PHP4 sessions > today, if you missed PHPLIB. Or go to Apple, have a look at > WebObjects. Or go to Microsoft, have a look at ASP+ and .NET and > their Visual Studio .NET if you still do not understand. I understand the developer-level hacks and kludges that have been created to help desktop coders transition to a web environment. But in many cases, these are simply tools to emulate some form of state management in an inherently stateless medium. The tools to create the exact same thing have been in PHP *since version 2*. What has been added is more libraries to create the appearance of traditional state, to abstract further from databases. But it's still just code to stash data somewhere until requested by a new connection, or code to reformat db interactions. Nothing magic, just storing data. > There is no difference between traditional applications and web > applications that cannot easily be hidden in an abstaction > layer. A pretty small abstraction layer even. But... are you having problems, or not? Apparently, you haven't been able to build a small abstraction layer to handle this issue, to emulate state and avoid re-execution? Or have you started with a small layer, which kept increasing in size as you added features? :-) > Once you can keep state (using PHPLIB or PHP 4 sessions), web > applications are just normal event driven applications like all > others, executed in an awkward and inefficient fashion. > Application servers are the means to end this awkwardness. PHP cannot end the stateless nature of the web. It can only emulate state in various ways. As you will discover on *all* "application servers", they are merely doing a similar thing, with similar hacks and kludges for different process models. They take data in, stash it into a db, a file, a memory cache. > > power the server, or change the design philosophy. On a > > desktop, you can assume the next action will be from the same > > user. > I am not talking desktop. I am talking middleware servers and > backoffice integration. 3000 concurrent users or open > transactions are not really a problem unless you do it backwards > and spend time compiling instead of comitting. Well, your goal then is to *reduce your compile times*. PHP can give you tools, but that can't repair inefficiencies in http data transactions.... this is why I've attempted to point you into directions to *reduce your compile times*, to eliminate object overhead, unused code loads, excessive abstractions. > > This is a security feature. The user can act as any given user, > > but if you are accepting internet connections over port 80, you > > have absolutely no way of determining who is, and isn't, actually > > sending them. > You are mixing several things up here. You are going nowhere as > long as you do this. > > We have > - user identities with in the application. > - user identities used by the application to authenticate itself > against a database. > - user identities used by the application to identitfy and > authenticate itself against the operating system (i.e. while > performing file operations). My point: You are basing this all on unknown traffic, on questionable data. A single apache id allows you to treat this the same way as you treat port 80... or are you suggesting that each user connect to some other port? Of course not. Apache is just a layer in between your web requests and your application. So, you can *create* the layer of authentication in your application, or use something simple like LDAP to manage it, and then code your transactions based on the access you've provided to those users. Using apache, you can create a safe, isolated, area which insulates you from multiple user ID's on the OS, even with only one user..... > In the Apache process model, running mod_php, your PHP code is > using arbitrary identities and authentication mechanisms in the > application. PHPLIB offers you any number of them. Yes, apply auth as the developers level. This is standard. Most middleware and RDBMS's come with auth that developers re-write anyways, because there are so many different appropriate models. Models which try to meet as many needs as possible become large, slow, and cumbersome. Apache allows you to apply auth on a number of db levels. PHP allows you to do this as well, without adding external code (though it does make is easier to use somebody else's, like PHPlib). Of course, if you are using a larger model, or a more multipurpoose model, you are weighed down by more code. > > How many > > active cursors can your DBMS handle? 128? Okay, if 5,000 users > > hit page one, and do nothing for the next 2 minutes, you now > > have 5,000 active db sessions, 5,000 cursors. Care to imagine > > the overhead of such a thing? > Have you ever heard the words "transaction monitor" or > "middleware"? Do you really think these are new problems which > have never been handled by anyone before? They handle it by programming for the web, by *not* actively maintaining state, or even making an effort to. Perhaps I'm digging deeper into the layers than you are normally comfortable with, but deep down, they're just reusing persistant connections, and storing a pseudo-"state" elsewhere. This is something you can do with PHP, if you want to write code for it. Store your data somewhere that your processes/threads can get to, cache it when loading, very simple. The multi-process model of apache makes this harder than a single-process model, but there are other options availabale. -Ronabop -- Personal: ron@opus1.com, 520-326-6109, http://www.opus1.com/ron/ Work: rchmara@pnsinc.com, 520-546-8993, http://www.pnsinc.com/ The opinions expressed in this email are not neccesarrily those of myself, my employers, or any of the other little voices in my head.

« previous php.dev (#40715) next »