Re: Large PHP/Database driven site & scalability
| From: | Eric Peters | Date: | Wed, 19 Jul 2000 15:30:15 +0000 |
| Subject: | Re: Large PHP/Database driven site & scalability | ||
| References: | 1 | Groups: | php.general |
| Request: | Send a blank email to php-general+get-7285@lists.php.net to get a copy of this message | ||
On Wed, 19 Jul 2000, Michael Kimsal wrote:
> We are in a similar situation with a particular client, but
> we may not end up using PHP (nothing to do with
> ability - it's client's 'comfort' with ASP, although they don't
> actually
> DO the coding, and imo the choice of language shouldn't be
> as important to them as the accuracy of the business logic, with
> the amount of money they have riding on the project).
right...we are in the boat of having a site already done in php and need
to redesign our connections/databases n stuff to handle an even larger
client base
>
> Anyway, we've looked at the pooling issue, and were starting
> to come up with a pool of Perl processes to proxy requests from
> PHP to a database. We couldn't effectively use the mysql_ functions
> so were looking at writing functions to open sockets on the local
> machine.
>
this is a good thought actually since you could specify a unix socket for
the mysql_pconnect to connect to - it would be interesting trying to write
a wrapper daemon process that would look for something like "sql
query:" in the queries and part of the syntax would include (what database
server?) or parsing off some use <database>; stuff
> Unfortunately, someone here formatted our dev machine, so
> our testcode is gone, but it wasn't working out that well anyway...
> (might have been for the best).
its definately a strategy I havn't thought of before...
>
> A new approach we're going to try is doing XMLRPC calls from PHP
> to a separate process, probably on a different machine. That
> XMLRPC server will keep track of connections, relay SQL calls,
> and send back the results. I'm not sure if the speed will be up to
> it, tho, and we may end up having to custom write something like this
> in C (unfortunately, none of us are C coders!).
XMLRPC?
I'm a C coder so its not that difficult finding a C coder - just more
difficult figuring out what needs to be done : )
>
> Not sure what you're driving at in #2. Setting up a separate
> server to proxy authentication requests to multiple db servers,
> depending on login data? (A-M = server 1, N-Z = server 2) ?
______ ______________ ______________________
| user | -> | www.blah.com | -> | blah.com says 'user' |
------ -------------- | is assigned to 'x' |
| server pool |
----------------------
|
_______ \/ _______
| www.x |------| www.y |
------- -------
\______/
|
\/
_________________
| db server with |
| info for 'user' |
-----------------
where there is a different webserver/database pool depending on the user
- a nice mod_rewrite rule would work well so the user would get forwarded
onto the webserver that is connected to the database with their shit on it
that help at all?
>
> We've not hit an issue yet on #3 where we had to split up
> data onto multiple SQL servers. Is there a reason you can't get
> bigger/faster drives and put it all on one machine? Performance?
>
> I realize this sounds theoretical, and it partially is, but we're
> working on 2 rather large clients - one has ~500 simultaneous
> users, with collectively around 5 million rows of data, but that
> DB is SQL7, and we haven't run into any limitations yet. Another
> semi large project is ~100 simultaenous users but less than
> 1 million rows of data (so far, but it's growing). This is MySQL
> on a dual 550 with 768 megs and a 3 drive SCSI RAID (or
> something like that).
we are logging a tremendous ammount of data - we are a y generation portal
site that has having issues with script kiddies out there decompiling java
applets and stuff to "cheat" the games (which they can trade their points
in for actual prizes) so from an auditing standpoint we are logging every
point earned and etc
we definately are starting to outgrow a single hefty database server
(dual p3 500 w/gig of ram and a 100 gig raid 5 array)
and are trying to identify if we can do this "with more brains than
balls" -
>
> Do you have the ability to make some tables HEAP tables,
> to keep them in RAM? This may speed up authentication queries.
> Also, how are you handling authentication? Hitting the
> db on every page? Our larger project has us do this, because
yeah there are several db calls per page - its based upon the phplib
session handling (though severly modd'd) but a cookie is still set that
includes the session and all the variables are set in a heap table for the
session data so that isn't really a big problem
its all of the other content that is making us hit limits - if we *just*
were able to syphen people off into a #2 environment it would be easy -
but we are trying to make it such a community driven site we are finding
we need to do searches across the site and relational queries between two
different users
> they don't want 2 people to use the same account at the same
> time, so we need to check for concurrent users. If you
> don't need to do that, a authentication can be verified by using
> some funky algorithmic stuff checking data in a cookie,
> rather than a db query every page. Sorry if that's elementary
that's already handled - definately not that big of an issue - we are
doing stuff cookie/session based
> to you - you are obviously putting a lot of thought into this
> planning ahead and have probably thought about many of these
> issues already.
>
> I hope some of this helps, and if you need any help testing things,
> we may have some spare cycles to donate.
Sounds good...definately would like info on how far you got with the perl
pooling and what major roadblocks you identified while trying to develop a
prototype - I think I identified the major issues I thought of(how it
would even know which server to talk to) - and I would be interested in
your response
>
> Thanks.
>
> Eric Peters wrote:
>
> > Basic problem:
> > distributed database system
> >
> > A little background info:
> >
> > The site is a community driven website that has over a quarter million
> > accounts, an average of 400 simultaneous users and well over 200 million
> > rows of data in various tables with logged information on points earned
> > during games, profile data, editorials, message postings and other misc
> > stored content
> >
> > We are running php4, mod_backhand, and MySQL 3.23.x
> >
> > We have so far been half addressing scalability with several hacks but
> > are seriously not having to sit down and readdress the issue
> > properly.
> >
> > We don't have the VC capital necessary to spend a couple hundred thousand
> > to a million dollars on an Oracle enterprise system - and need to address
> > the situation with more brains than balls
> >
> > With several senarios we have drawn for linking web servers to database
> > servers, with a myriad of "gateway" database servers setup for
> > distributing user data and linking webservers to the appropriate db server
> > that stores profile information there are a couple key issues I would love
> > to have people comment on
> >
> > 1) how can you get a mysql_pconnect to use an apache server connection
> > pool (rather than a child persistant connection) - yes I know its not
> > implemented - but how hard would it be - and how much would it cost to get
> > someone to do it : ) - and from usage standpoint of a database driven
> > website is it practical to have a pool of 50 on a server with 255 httpd
> > children ? assuming every page had at *least* one database query..
> >
> > 2) looking at an authentication server that would have a
> > username/password/database assignment - would it be practical to stepup to
> > having a "datapipe" that would automatically query an appropriate sql
> > server and retrieve the information - and how hard would a fd limit be
> > reached etc
> >
> > 3) in a user->database server assignment senario - what about information
> > that needs to be shared - or searched across the database - say for some
> > "find a friend" where it matches people with people - or game ranking
> > system where you want to list all of the top 20 players or a myrad of
> > other options that would require a plethora of queries across each server
> > - is there a particular approach people find themselves using - whether
> > there is a single server that holds "current" data and it then would be
> > archived to an older server - how reliable is it and etc
> >
> > in general i'm looking for applied knowledge that people have tried,
> > rather than theoretical knowledge - i am definately interested in finding
> > a niche of peers that have to deel with the same issues and create a
> > support network
> >
> > If i'm way off base let me know,
> >
> > I appreciate your time and thank you,
> >
> > Eric
> >
> > --
> > PHP General Mailing List (http://www.php.net/)
> > To unsubscribe, e-mail: php-general-unsubscribe@lists.php.net
> > For additional commands, e-mail: php-general-help@lists.php.net
> > To contact the list administrators, e-mail: php-list-admin@lists.php.net
>
> --
> ==========================
> Michael Kimsal
> http://www.tapinternet.com
> 734-480-9961
>
>
>