TiCA: A Tibetan Text Compression Algorithm

Uložené v:
Podrobná bibliografia
Názov: TiCA: A Tibetan Text Compression Algorithm
Autori: Chen Shuo, Suonan Jiancuo, Nima Zhaxi, Renqing Nuobu
Zdroj: Proceedings of the 2nd International Conference on Artificial Intelligence and Advanced Manufacture. :8-12
Informácie o vydavateľovi: ACM, 2020.
Rok vydania: 2020
Predmety: 4. Education, 0202 electrical engineering, electronic engineering, information engineering, 02 engineering and technology
Popis: This paper proposes a Tibetan text compression algorithm (TiCA), which is based on the fact that each Tibetan syllable is composed of one to seven components and each component has a unique Unicode encoding. First of all, through statistical analysis of 20G Tibetan text corpus, a fault-tolerant mapping dictionary is established and used as the dictionary of the TiCA. The TiCA then compresses the Tibetan text according to the mapping dictionary by mapping the original code to a single code. Finally, the experimental comparison shows that the Tibetan text compression algorithm proposed in this paper has achieved excellent results both in the compression rate and time consuming.
Druh dokumentu: Article
DOI: 10.1145/3421766.3421868
Prístupová URL adresa: https://dblp.uni-trier.de/db/conf/aiam/aiam2020.html#SuonanCRZ20
https://doi.org/10.1145/3421766.3421868
Rights: URL: https://www.acm.org/publications/policies/copyright_policy#Background
Prístupové číslo: edsair.doi.dedup.....665d19c20e9fb8cc734e86d18f7e38a5
Databáza: OpenAIRE
Popis
Abstrakt:This paper proposes a Tibetan text compression algorithm (TiCA), which is based on the fact that each Tibetan syllable is composed of one to seven components and each component has a unique Unicode encoding. First of all, through statistical analysis of 20G Tibetan text corpus, a fault-tolerant mapping dictionary is established and used as the dictionary of the TiCA. The TiCA then compresses the Tibetan text according to the mapping dictionary by mapping the original code to a single code. Finally, the experimental comparison shows that the Tibetan text compression algorithm proposed in this paper has achieved excellent results both in the compression rate and time consuming.
DOI:10.1145/3421766.3421868