A Proposed UNICODE-Based Extended Romanization System for Persian Texts

So far, various Romanization schemes have been proposed for capturing Persian text using Latin alphabet. However, each have served a very specific and yet limited function. This paper proposes an extended Romanization scheme that can facilitate a wide range of encoding needed in the field of Natural...

Full description

Saved in:
Bibliographic Details
Published in:International journal of information science and management Vol. 10; no. 1; pp. 57 - 71
Main Author: M. A. Mahdavi
Format: Journal Article
Language:English
Published: Regional Information Center for Science and Technology (RICeST) 01-07-2012
Online Access:Get full text
Tags: Add Tag
No Tags, Be the first to tag this record!
Description
Summary:So far, various Romanization schemes have been proposed for capturing Persian text using Latin alphabet. However, each have served a very specific and yet limited function. This paper proposes an extended Romanization scheme that can facilitate a wide range of encoding needed in the field of Natural Language Processing. The proposed scheme endeavors to preserve both orthographic and phonological phenomena in the language. It also accounts for encoding handwritten manuscripts, in which glyph ambiguity is a salient feature. It is particularly relevant to Romanizing the Kufi script, in which diacritical marks are omitted. The current work also recommends orthographic rules in an effort to standardize future Romanization tasks.
ISSN:2008-8302
2008-8310