原理:
如果***type(text) is bytes***,
那么text.decode('unicode_escape')
如果type(text) is str,
那么text.encode(‘latin1’).decode(‘unicode_escape’)
1. 案例:
*
#coding=utf-8
import requests,re,json,traceback
from bs4 import BeautifulSoup
def qiushibaike():
content = requests.get('http://baike.baidu.com/city/api/citylemmalist?type=0&cityId=360&offset=1&limit=60').content
soup = BeautifulSoup(content, 'html.parser')
print(soup.prettify()) #.decode("unicode_escape")
#目前soup.prettify()为str
new=soup.prettify().encode('latin-1').decode('unicode_escape')
#.dencode('latin-1').encode('latin-1').decode('unicode_escape')
print(new)
if __name__=='__main__':
qiushibaike()
2. 结果对比:
另外爬取时,网站代码出现GBK无法编译python3,如出现如下:
ÖйúÉÙÊýÃñ×åÌØÉ«´åÕ¯[6]
示例:
#coding=utf-8
import requests
#共有6页,首页为空不为6
for i in range(6):
if i==0:
url='http://www.tcmap.com.cn/list/zhongguoshaoshuminzutesecunzhai.html'
else:
url='http://www.tcmap.com.cn/list/zhongguoshaoshuminzutesecunzhai'+str(i)+'.html'
response=requests.get(url)
print(type(response))
#如需成功编译,在.TEXT下面增加#号部分
html=response.text #.encode('latin-1').decode('GBK')
print(html)