Python中中文路径处理问题的研究

最新推荐文章于 2023-09-10 15:05:07 发布

原创最新推荐文章于 2023-09-10 15:05:07 发布 · 3.1k 阅读

0 ·

CC 4.0 BY-SA版权

文章标签：

#python #编码 #中文 #路径 #utf-8

原创同时被 3 个专栏收录

7 篇文章

订阅专栏

Python

7 篇文章

订阅专栏

总结

3 篇文章

订阅专栏

本文详细解析了Python中Unicode编码的应用与处理方法，包括字符串拼接、路径处理及不同编码下的问题解决策略。

部署运行你感兴趣的模型镜像

a = '你' 为 str 对象
a = u'你' 为 unicode 对象

>>> print 'u' + '你'
>>> u浣
输出乱码

>>> print 'u' + u'你'
>>> u你
正常

>>> print 'u你'
>>> u浣

输出乱码

4.
>>> print 'u你' + 'u'
>>> u浣爑
输出乱码

>>> print u'u你' + 'u'
>>> u你u
正常

>>> print u'u你' + '你'
出现错误 UnicodeDecodeError: 'ascii' codec can't decode byte 0xe4 in position 0: ordinal not in range(128)
分析：'你'在内存中为 0xe4，而python默认的编码方案是ascii，ascii无法识别0xe4

>>> print u'u你' + u'你'
>>> u你你
正常

8.
>>> print 'u你' + u'你'
出现错误 UnicodeDecodeError: 'ascii' codec can't decode byte 0xe4 in position 1: ordinal not in range(128)

9.
>>> print 'u你'.decode('utf-8') + u'你'
>>> u你你
正常

10.
而在处理由系统采集的含有中文的路径时，使用string.decode('utf-8')就不一定行了，因为简体中文的windows系统默认编码为gb2312，繁体中文版会采用Big5码
实验过程如下：
file_from = sys.argv[1] 为由系统采集的包含中文的路径
file_to = file_from[:file_from.rfind('\\')+1].decode('utf-8') + u'你_' + file_from[file_from.rfind('\\')+1:].decode('utf-8')

print file_to

将出现错误：UnicodeDecodeError: 'utf8' codec can't decode byte 0xbb in position 24: invalid start byte

应该使用：decode('gb2312')
file_to = file_from[:file_from.rfind('\\')+1].decode('gb2312') + u'你_' + file_from[file_from.rfind('\\')+1:].decode('gb2312')

print file_to 正常

11.

而如果file_from是由你自己写入的包含中文的路径，如file_from = ‘c:\你.txt’

那么就应该用decode('utf-8')

可以参考上面的第7点和第9点

不足及错误之处，请批评指正！！谢谢！！

参考文章：

Why you benefit from using UTF-8 Unicode everywhere in your web applications

Python "'ascii' codec can't decode byte" explained and how to solve it

Windows 记事本的 ANSI、Unicode、UTF-8 这三种编码模式有什么区别？

您可能感兴趣的与本文相关的镜像

AutoGPT

AI应用

AutoGPT于2023年3月30日由游戏公司Significant Gravitas Ltd.的创始人Toran Bruce Richards发布,AutoGPT是一个AI agent（智能体），也是开源的应用程序，结合了GPT-4和GPT-3.5技术，给定自然语言的目标，它将尝试通过将其分解成子任务，并在自动循环中使用互联网和其他工具来实现这一目标