more Unicode data updates

Previous Topic Next Topic
 
classic Classic list List threaded Threaded
3 messages Options
Reply | Threaded
Open this post in threaded view
|

more Unicode data updates

Peter Eisentraut-6
src/include/common/unicode_norm_table.h also should be updated to the
latest Unicode tables, as described in src/common/unicode.  See attached
patches.  This also passes the tests described in
src/common/unicode/README.  (That is, the old code does not pass the
current Unicode test file, but the updated code does pass it.)

I also checked contrib/unaccent/ but it seems up to date.

It seems to me that we ought to make this part of the standard major
release preparations.  There is a new Unicode standard approximately
once a year; see <https://unicode.org/Public/>.  (The 13.0.0 listed
there is not released yet.)

It would also be nice to unify and automate all these "update to latest
Unicode" steps.

--
Peter Eisentraut              http://www.2ndQuadrant.com/
PostgreSQL Development, 24x7 Support, Remote DBA, Training & Services

0001-Correct-script-name-in-README-file.patch (1K) Download Attachment
0002-Make-script-output-more-pgindent-compatible.patch (1K) Download Attachment
0003-Update-unicode_norm_table.h-to-Unicode-12.1.0.patch (251K) Download Attachment
Reply | Threaded
Open this post in threaded view
|

Re: more Unicode data updates

Thomas Munro-5
On Thu, Jun 20, 2019 at 8:35 AM Peter Eisentraut
<[hidden email]> wrote:

> src/include/common/unicode_norm_table.h also should be updated to the
> latest Unicode tables, as described in src/common/unicode.  See attached
> patches.  This also passes the tests described in
> src/common/unicode/README.  (That is, the old code does not pass the
> current Unicode test file, but the updated code does pass it.)
>
> I also checked contrib/unaccent/ but it seems up to date.
>
> It seems to me that we ought to make this part of the standard major
> release preparations.  There is a new Unicode standard approximately
> once a year; see <https://unicode.org/Public/>.  (The 13.0.0 listed
> there is not released yet.)
>
> It would also be nice to unify and automate all these "update to latest
> Unicode" steps.

+1, great idea.  Every piece of the system that derives from Unicode
data should derive from the same version, and the version should be
mentioned in the release notes when it changes, and should be
documented somewhere centrally.  I wondered about that when working on
the unaccent generator script but didn't wonder hard enough.

--
Thomas Munro
https://enterprisedb.com


Reply | Threaded
Open this post in threaded view
|

Re: more Unicode data updates

Peter Eisentraut-6
In reply to this post by Peter Eisentraut-6
On 2019-06-19 22:34, Peter Eisentraut wrote:
> src/include/common/unicode_norm_table.h also should be updated to the
> latest Unicode tables, as described in src/common/unicode.  See attached
> patches.  This also passes the tests described in
> src/common/unicode/README.  (That is, the old code does not pass the
> current Unicode test file, but the updated code does pass it.)

committed

--
Peter Eisentraut              http://www.2ndQuadrant.com/
PostgreSQL Development, 24x7 Support, Remote DBA, Training & Services