A Proposed UNICODE-Based Extended Romanization System for Persian Texts
So far, various Romanization schemes have been proposed for capturing Persian text using Latin alphabet. However, each have served a very specific and yet limited function. This paper proposes an extended Romanization scheme that can facilitate a wide range of encoding needed in the field of Natural...
Saved in:
Published in: | International journal of information science and management Vol. 10; no. 1; pp. 57 - 71 |
---|---|
Main Author: | |
Format: | Journal Article |
Language: | English |
Published: |
Regional Information Center for Science and Technology (RICeST)
01-07-2012
|
Online Access: | Get full text |
Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Summary: | So far, various Romanization schemes have been proposed for capturing Persian text using Latin alphabet. However, each have served a very specific and yet limited function. This paper proposes an extended Romanization scheme that can facilitate a wide range of encoding needed in the field of Natural Language Processing. The proposed scheme endeavors to preserve both orthographic and phonological phenomena in the language. It also accounts for encoding handwritten manuscripts, in which glyph ambiguity is a salient feature. It is particularly relevant to Romanizing the Kufi script, in which diacritical marks are omitted. The current work also recommends orthographic rules in an effort to standardize future Romanization tasks. |
---|---|
ISSN: | 2008-8302 2008-8310 |