Re: cvs: phpdoc / Makefile.in
| From: | Friedhelm Betz | Date: | Sat, 17 May 2003 16:19:19 +0000 |
| Subject: | Re: cvs: phpdoc / Makefile.in | ||
| References: | 1 2 3 | Groups: | php.doc |
| Request: | Send a blank email to phpdoc+get-969353572@lists.php.net to get a copy of this message | ||
On Friday 16 May 2003 15:56, Gabor Hojtsy wrote:
[...]
> Hm, AFAIK xsltproc works with iconv, so probably only the iconv
> supported encodings can be used. The windows version requires the iconv
> dll to be in place, but I cannot find any dependencies related to iconv
> on linux. Though there is probably a dependency. The question is up for
> the doc-he guys, what encoding they are able to work with?
Qoute from xmlsoft.org (http://xmlsoft.org/encoding.html):
<qoute>
Default supported encodings
libxml has a set of default converters for the following encodings
(located in encoding.c):
1. UTF-8 is supported by default (null handlers)
2. UTF-16, both little and big endian
3. ISO-Latin-1 (ISO-8859-1) covering most western languages
4. ASCII, useful mostly for saving
5. HTML, a specific handler for the conversion of UTF-8 to ASCII with
HTML predefined entities like © for the Copyright sign.
More over when compiled on an Unix platform with iconv support the full
set of encodings supported by iconv can be instantly be used by libxml.
On a linux machine with glibc-2.1 the list of supported encodings and
aliases fill 3 full pages, and include UCS-4, the full set of
ISO-Latin encodings, and the various Japanese ones.
</qoute>
To get supported encodings by iconv try iconv -l and iso-8859-8-i is not among
them, iso-8859-8 and WINDOWS-1255 are.
Besides that, encoding should be consistent in the translated he files,
currently there is:
<?xml version="1.0" encoding="iso-8859-8-i"?>
<?xml version="1.0" encoding="iso-8859-8"?>
<?xml version="1.0" encoding="iso-8859-1"?>
AFAIK the input encoding should be 8859-8 or WINDOWS-1255. Output encoding is
set to 8859-8 by the xsl-sheets. The question: why is iso-8859-8-i used in he
files?
The entiity files:
language-defs and language-snippets might be changed, as suggested by shimi
iconv -f windows-1255 -t utf-8 to make xsltproc happy;-)
At least my tests have shown, that this works fine, at least xmllint and
xsltproc doesn't complain about encodings. Can't say anything about the
results, because i can't read hebrew:-)
Testing:
1. I applied iconv conversion to the two entity files
2. changed encoding to WINDOWS-1255 in the curl folder of he tree and other
files in question (maybe this should be 8859-8).
Results (processed with xsltproc) can be found online at:
www.holliwell.de/he/
Friedhelm
p.s.: changing the language entitiy files has the effect, that make test
reports errors about invalid sgml characters.