Re: Script to Fetch Elephpants
| From: | Hannes Magnusson | Date: | Sun, 18 Sep 2011 16:46:33 +0000 |
| Subject: | Re: Script to Fetch Elephpants | ||
| References: | 1 2 3 4 5 | Groups: | php.webmaster |
| Request: | Send a blank email to php-webmaster+get-12213@lists.php.net to get a copy of this message | ||
On Sun, Sep 18, 2011 at 17:45, Stewart Lord <stewey@ambitious.ca> wrote:
> On 2011-09-18, at 3:37 AM, Hannes Magnusson <hannes.magnusson@gmail.com> wrote:
>
>> On Sun, Sep 18, 2011 at 07:17, Stewart Lord <stewey@ambitious.ca> wrote:
>>>
>>> On 2011-09-14, at 11:57 PM, Hannes Magnusson wrote:
>>>
>>>> Is url_sq guaranteed never to contain urlencoded paths, or
>>>> non-windows-friendly-filenames?
>>>> I doubt we have any Windows mirrors though.. but who knows.
>>>>
>>>> I notice you never cleanup the fetched images, so we will quickly have
>>>> bucketload of images.. maybe thats the idea - then we can shuffle them
>>>> on the mirrors?
>>>>
>>>> Its also missing sanitychecks around the file_get_contents() for the
>>>> json data, and ensuring $decoded actually contains anything before
>>>> overwriting photos.json.
>>>> Same when fetching the file, it needs to verify it actually fetched something.
>>>> ..And maybe check if the file already exists before overwriting it?
>>>>
>>>> -Hannes
>>>
>>>
>>> Hi Hannes,
>>>
>>> Thanks again for the code review. I've updated the script to incorporate your
>>> feedback:
>>>
>>> http://pastebin.com/BBu39Rv4
>>>
>>> The url_sq basename seems to only ever contain alpha-numerics, '_' and
>>> '.', but I've added a filter to make sure. I added a bit of logic at the end of the
>>> script to remove any stale images (as long as we set the limit reasonably high we should have enough
>>> to shuffle through). Also added sanity checks.
>>>
>>
>>
>> Looks good.
>> Is the result set random, or the "latest images"?
>> Just wondering how often new pics would be downloaded when we run this
>> every hour.
>>
>> -Hannes
>
> The API docs don't actually say, but it is returning the newest images. No need to run it
> every hour, but no harm either. It will just be the one service call, plus one (small) download per
> new image.
>
> I'm shuffling them when they are displayed. Not sure how many we want the script to fetch.
> I would say at least 100 and flickr places an upper limit of 400 per API call.
>
100 seems fair, then we have several to shuffle through and make to
frontpage seem less static.
The update-backend script is executed via cronjob once an our, so if
you include this there.. No need to do some extra time checks in it,
or add new cronjobs imo.
-Hannes