Introduction to text analysis

Computer Mashup

 

The amount of textual data available to us grows each day, comprising a huge resource which is potentially of tremendous value. This workshop will introduce the process of extracting meaningful structured data and information from text.  

Participants will use Python and NLTK (Natural Language Toolkit) for the retrieval of basic textual information such as:  

  • Word frequencies  

  • Plots of frequency distributions  

  • Common word pairs  

  • Part of Speech tagging 

The workshop will take place on Microsoft Teams via the Silent Disco mode. This is an asynchronous event, this means that when attending, you will work through the tutorial at your own pace with an instructor available online to help you with any issues. 

This is an intermediate level event. Intermediate events explore specific aspects of the method (libraries, tools etc.) and offer more in-depth understanding of the course topics, without introducing the basics. Some previous knowledge is required to be able to follow the content. 

Those who have registered to take part will receive an email with full details and a link to join the session in advance of the start time. 

After taking part in this event, you may decide that you need some further help in applying what you have learnt to your research. If so, you can book a Data Surgery meeting with one of our training fellows. 

More details about Data Surgeries. 

If you’re new to this training event format, or to CDCS training events in general, read more on what to expect from CDCS training. Here you will also find details of our cancellation and no-show policy, which applies to this event. 

Return to the Training Homepage to see other available events. 

You might be interested in

Graphic for a workshop titled ‘Text Classification in Practice: From Topic Models to Transformers.’ The background shows handwritten historical letters. A large green ampersand featuring an illustration of Ada Lovelace is placed on the left. The logo of the Centre for Data, Culture & Society (DCS) appears in the top right corner.

Text Classification in Practice: From Topic Models to Transformers

Graphic for a workshop titled ‘Using Prompting Efficiently for Research.’ The background shows an aged, torn book page with visible text. A large green ampersand featuring an illustration of Ada Lovelace is placed on the left. The logo of the Centre for Data, Culture & Society (DCS) appears in the top right corner.

Using Prompting Efficiently for Research

Graphic for a workshop titled ‘Foundations of Sentiment Analysis.’ The background is a sepia photograph of people working at desks in a large hall with overhead lamps. A large green ampersand featuring an illustration of Ada Lovelace is placed on the left. The logo of the Centre for Data, Culture & Society (DCS) appears in the top right corner.

Silent Disco: Foundations of Sentiment Analysis

Graphic for a workshop titled ‘Working Collaboratively Through Version Control.’ The background is a black-and-white photograph of people weaving on large looms. A large magenta ampersand featuring an illustration of Ada Lovelace is placed on the left. The logo of the Centre for Data, Culture & Society (DCS) appears in the top right corner.

Working Collaboratively through Version Control

Graphic for a workshop titled ‘Foundations of Webscraping.’ The background is a black-and-white photograph of students working together in a design studio with maps and models. A large teal ampersand featuring an illustration of Ada Lovelace is placed on the left. The logo of the Centre for Data, Culture & Society (DCS) appears in the top right corner.

Collecting Data from the Web: Foundation of Webscraping

Graphic for a workshop titled ‘Using API for Research.’ The background is a black-and-white photograph of people working with printing equipment and patterned sheets. A large magenta ampersand featuring an illustration of Ada Lovelace is placed on the left. The logo of the Centre for Data, Culture & Society (DCS) appears in the top right corner.

Using API for Research

Graphic for a workshop titled ‘Introduction to Geographical Data with QGIS.’ The background shows an old map of the world with detailed illustrations. A large teal ampersand featuring an illustration of Ada Lovelace is placed on the left. The logo of the Centre for Data, Culture & Society (DCS) appears in the top right corner.

Intro to Geographical Data with QGIS

Graphic for a workshop titled ‘Getting Started with Inferential Statistics.’ The background is a black-and-white photograph of people studying in a library with partitioned desks. A large teal ampersand featuring an illustration of Ada Lovelace is placed on the left. The logo of the Centre for Data, Culture & Society (DCS) appears in the top right corner.

Getting Started with Inferential Statistics