Scraping Websites with R

Book Now
Book now
Scraping websites with R

 

In Person 

This intermediate workshop will teach you how to scrape user-generated content from the internet using R. The workshop will start with a theoretical introduction to web scraping and specific approaches to scraping static websites with a focus on HTML tags. We will then select a website and go through the process of scraping it. Then we will discuss different kinds of websites that present problems for web scraping.  

This is an intermediate-level course. Students must have a basic background in R. This includes, at least the basic data types in R, how to install and load packages, and how to use functions, pipes, and apply/map functions. It will be sufficient for students to have taken the Introduction to Programming with R and RStudio course. 

Those who have registered to take part will receive an email with full details on how to get ready for the workshop. 

 

This workshop will be taught by Jessica Witte. 

After taking part in this event, you may decide that you need some further help in applying what you have learnt to your research. If so, you can book a Data Surgery meeting with one of our training fellows. 

More details about Data Surgeries. 

If you’re new to this training event format, or to CDCS training events in general, read more on what to expect from CDCS training. Here you will also find details of our cancellation and no-show policy, which applies to this event. 

This training is connected to the Legal and Ethical issues in Webscraping training taking place on the 14th of October.   

 

If you're interested in other training on webscraping and text analysis, you can have a look at the following: 

 

 

Return to the Training Homepage to see other available events. 

Room 4.35, Edinburgh Futures Institute

This room is on Level 4, in the North East side of the building.

When you enter via the level 2 East entrance on Middle Meadow Walk, the room will be on the 4th floor straight ahead.

When you enter via the level 2 North entrance on Lauriston Place underneath the clock tower, the room will be on the 4th floor to your left.

When you enter via the level 0 South entrance on Porters Walk (opposite Tribe Yoga), the room will be on the 4th floor to your right.

You might be interested in

Graphic for an event titled ‘BYOD Festival.’ The background is a black-and-white photograph of people sitting around a table, drinking tea and playing cards. A large magenta ampersand featuring an illustration of Ada Lovelace is placed on the left. The logo of the Centre for Data, Culture & Society (DCS) appears in the top right corner.

Bring Your Own Data (BYOD) Fest

Graphic for a workshop titled ‘Using API for Research.’ The background is a black-and-white photograph of people working with printing equipment and patterned sheets. A large magenta ampersand featuring an illustration of Ada Lovelace is placed on the left. The logo of the Centre for Data, Culture & Society (DCS) appears in the top right corner.

Using API for Research

a yellow tinged photo of people entering a building, with the text "Brad Rittenhouse, Project Deep Dive"

Giving Humanists a Helping Hand in HPC

Graphic for a workshop titled ‘Text Classification in Practice: From Topic Models to Transformers.’ The background shows handwritten historical letters. A large green ampersand featuring an illustration of Ada Lovelace is placed on the left. The logo of the Centre for Data, Culture & Society (DCS) appears in the top right corner.

Text Classification in Practice: From Topic Models to Transformers

Graphic for a workshop titled ‘Getting Started with Descriptive Statistics.’ The background is a black-and-white photograph of people reading and working in a library. A large magenta ampersand featuring an illustration of Ada Lovelace is placed on the left. The logo of the Centre for Data, Culture & Society (DCS) appears in the top right corner.

Getting Started with Descriptive Statistics

Graphic for a workshop titled ‘Foundations of Webscraping.’ The background is a black-and-white photograph of students working together in a design studio with maps and models. A large teal ampersand featuring an illustration of Ada Lovelace is placed on the left. The logo of the Centre for Data, Culture & Society (DCS) appears in the top right corner.

Collecting Data from the Web: Foundation of Webscraping

an old map of Acotland with the text "Jennifer Smith & Brian Aitken, Project deep Dive"

Who Speaks Scots Where: What Crowdsourcing Reveals

Graphic for a workshop titled ‘Using Prompting Efficiently for Research.’ The background shows an aged, torn book page with visible text. A large green ampersand featuring an illustration of Ada Lovelace is placed on the left. The logo of the Centre for Data, Culture & Society (DCS) appears in the top right corner.

Using Prompting Efficiently for Research