Silent Disco: Cleaning Data with OpenRefine

Silent Disco mashup

 

Online

Our 'Silent Disco' workshops are based on tutorials from the Programming Historian website. This training event will follow content from the tutorial, Cleaning Data with OpenRefine.   

Open Refine is a powerful tool for working with messy data: cleaning it; transforming it from one format into another  

In this lesson, you will learn the principles and practice of data cleaning, as well as how OpenRefine can be used to perform four essential tasks that will help you to clean your data:  

  1. Remove duplicate records  

  1. Separate multiple values contained in the same field  

  1. Analyse the distribution of values throughout a data set  

  1. Group together different representations of the same reality  

This is an asynchronous event: when taking part, you will work through the tutorial at your own pace with an instructor available online to help you with any issues.  

Participants will meet for a brief introduction to OpenRefine, and will then work on the tutorial at their own pace. The facilitator will be available via Teams Chat to reply to any questions that arise during the workshop.  

This is a beginner level workshop. No previous knowledge on the topic is required/expected and the trainer will cover the basics of the method.  

To attend this course, you will have to join the associated Microsoft Teams group. The link to join the group will be sent to the attendees prior to the course start date, so please make sure to do so in advance.  

After taking part in this event, you may decide that you need some further help in applying what you have learnt to your research. If so, you can book a Data Surgery meeting with one of our training fellows.  

More details about Data Surgeries.  

If you’re new to this training event format, or to CDCS training events in general, read more on what to expect from CDCS training. Here you will also find details of our cancellation and no-show policy, which applies to this event.  

If you're interested in other training on Data Wrangling, have a look at the following:

 

Return to the Training Homepage to see other available events.

You might be interested in

Graphic for an event titled ‘BYOD Festival.’ The background is a black-and-white photograph of people sitting around a table, drinking tea and playing cards. A large magenta ampersand featuring an illustration of Ada Lovelace is placed on the left. The logo of the Centre for Data, Culture & Society (DCS) appears in the top right corner.

Bring Your Own Data (BYOD) Fest

Graphic for a workshop titled ‘Working Collaboratively Through Version Control.’ The background is a black-and-white photograph of people weaving on large looms. A large magenta ampersand featuring an illustration of Ada Lovelace is placed on the left. The logo of the Centre for Data, Culture & Society (DCS) appears in the top right corner.

Working Collaboratively through Version Control

Graphic for a workshop titled ‘Text Classification in Practice: From Topic Models to Transformers.’ The background shows handwritten historical letters. A large green ampersand featuring an illustration of Ada Lovelace is placed on the left. The logo of the Centre for Data, Culture & Society (DCS) appears in the top right corner.

Text Classification in Practice: From Topic Models to Transformers

Graphic for a workshop titled ‘Using Prompting Efficiently for Research.’ The background shows an aged, torn book page with visible text. A large green ampersand featuring an illustration of Ada Lovelace is placed on the left. The logo of the Centre for Data, Culture & Society (DCS) appears in the top right corner.

Using Prompting Efficiently for Research

Graphic for a workshop titled ‘Foundations of Sentiment Analysis.’ The background is a sepia photograph of people working at desks in a large hall with overhead lamps. A large green ampersand featuring an illustration of Ada Lovelace is placed on the left. The logo of the Centre for Data, Culture & Society (DCS) appears in the top right corner.

Silent Disco: Foundations of Sentiment Analysis

Graphic for a workshop titled ‘Introduction to Geographical Data with QGIS.’ The background shows an old map of the world with detailed illustrations. A large teal ampersand featuring an illustration of Ada Lovelace is placed on the left. The logo of the Centre for Data, Culture & Society (DCS) appears in the top right corner.

Intro to Geographical Data with QGIS

an old map of Acotland with the text "Jennifer Smith & Brian Aitken, Project deep Dive"

Who Speaks Scots Where: What Crowdsourcing Reveals