Re[4]: [PEAR-DEV] Package proposal
| From: | Maxim Antipin | Date: | Tue, 12 Oct 2004 13:04:47 +0000 |
| Subject: | Re[4]: [PEAR-DEV] Package proposal | ||
| References: | 1 2 | Groups: | php.pear.dev |
| Request: | Send a blank email to pear-dev+get-33798@lists.php.net to get a copy of this message | ||
Hello Kouber,
Tuesday, October 12, 2004, 6:21:08 PM, you wrote:
KS> I have a couple of objections both on the sense of such a package and on
KS> your implementation.
KS> 1. Imagine a situation when 2 clients on 2 different machines (so, 2
KS> different phps) are making requests to one and the same database through
KS> this caching package. You will have 2 different caches - so 2 wrong "views"
KS> of the database, because the first client couldn't know what changes are
KS> made by the second one and vice versa.
No, you are wrong. Caches are located not on clients' machines. There
are no several copies of cache. All clients use one cache. The only
possibility to get error if MySQL will receive two queries
'simultaneously'. One should modify data second should query. Then if
they will executed with interval less the 1ms (average time needed to
process modification query and update cache storage) then we will get
rotten data. But it is problem of all disk caches.
KS> 2. Many of the hosting companies provide some database administration tool,
KS> such as phpMyAdmin for example. As you have to agree, you can't force
KS> phpMyAdmin to use your caching mechanism. Even if it's possible, it will
KS> create again one "general cache" for each user - the main purpose you
KS> mentioned for not using MySQL's native caching
Yes i can not. But why it bother you :) Common script user usually do
not modifies data directly through phpMyAdmin. It just sets up
application and runs it. If you are not common user and you modified
data bypassing caching mechanism you can just delete cache from disk.
That is all you need to invalidate all cache :)
KS> 3. Some of the hosting companies provide also ssh access - so the user is
KS> free to use the mysql's shell tool.
See above.
KS> 4. When the user imports a db dump via: mysql < dump.sql, or with some other
KS> tool - the cache will stay untouched and inconsistent.
See above.
KS> 5. I can't believe that it will be faster to go through the cache file and
KS> search for the query each time, than just to use MySQL's native caching. I'm
KS> not a hosting company, but it sounds very strange for me to provide old
KS> MySQL (3.x) and new PHP (with PEAR and with that package included).
Yes it is strange and it is very uncomfortable but it real.
KS> I had no time to look at every piece of code, but I've found already some
KS> bug-potential things:
KS> - In processQuery($query) method you assume that the query begins with the
KS> SQL command - no spaces or comments. See
KS> http://dev.mysql.com/doc/mysql/en/Query_Cache_How.html for an
KS> example:
KS> "Before MySQL 5.0, a query that begins with a leading comment might be
KS> cached, but could not be fetched from the cache. This problem is fixed in
KS> MySQL 5.0."
KS> In the same check you have to include also TRUNCATE and DROP commands.
Yes i know about this. And it can be fixed easily. Now we discuss
wether PHP community is interested in MySQL driver with caching.
KS> - Maybe I'm wrong, but it seems that you invoke getTables() each time you
KS> create a new instance of the object, which executes mysql_list_tables(),
KS> which is of course sent to the server...so you actually have the network
KS> overhead before any cache check is possible.
You are wrong. Function 'mysql_list_tables' is called in cache
constructor. But cache constructor is called once when we create DB
driver. When driver processes queries it does not call
'mysql_list_tables' function.
KS> - In _tablesInQuery($query) method you are using strpos() to check if a
KS> table is in a query or not. Now imagine that you have aliases or functions
KS> or just strings with the name of the table - your check will return true,
KS> even when the table doesn't exist in the query. You have to use some regular
KS> expressions here to determine the tables involved, rather than retrieving
KS> all the tables once, and looking for these *strings* in the query.
:)) Yes, idea is good. To be more specific, i need use not regexps (As
they allow describe only language with finite state grammar) but
context-restricted grammar analyzer. But the main point of cache is
speed, so i implemented most simple and fast way. We can treat it as
limitation: "All table names should be unique among other SQL
entities".
KS> Anyway, I think the first few points I just mentioned in the beginning are
KS> very important - the inconsistency risk here is just too high and I strongly
KS> believe that database caching should be done on the database.
Well, it is your opinion and i will not try to convince you to change
it. Just read my previous post where i explained disadvantages of
current implementation of MySQL internal cache.
I want to say that this is not ideal caching solution and it has some
restrictions and disadvantages. But it can be useful, at least it is
useful for us :)
--
Best regards,
Maxim Antipin (ITScript CEO) mailto:max@itscript.com