Bonnes pratiques en documentation
46.4K views | +6 today
Follow
 
Scooped by Stéphane Cottin
onto Bonnes pratiques en documentation
Scoop.it!

The Code4Lib Journal – Approaching the largest ‘API’: extracting information from the Internet with Python

The Code4Lib Journal – Approaching the largest ‘API’: extracting information from the Internet with Python | Bonnes pratiques en documentation | Scoop.it
This article explores the need for libraries to algorithmically access and manipulate the world’s largest API: the Internet. The billions of pages on the ‘Internet API’ (HTTP, HTML, CSS, XPath, DOM, etc.) are easily accessible and manipulable. Libraries can assist in creating meaning through the datafication of information on the world wide web. Because most information is created for human consumption, some programming is required for automated extraction. Python is an easy-to-learn programming language with extensive packages and community support for web page automation. Four packages (Urllib, Selenium, BeautifulSoup, Scrapy) in Python can automate almost any web page for all sized projects. An example warrant data project is explained to illustrate how well Python packages can manipulate web pages to create meaning through assembling custom datasets.
more...
No comment yet.
Bonnes pratiques en documentation
Dernieres informations sur les bonnes pratiques en recherche documentaire, analyse de la documentation, moteurs de recherche,...
Curated by Stéphane Cottin