1
00:00:00,980 --> 00:00:01,970
Hey, welcome back.

2
00:00:02,000 --> 00:00:09,200
Yesterday, we did some analysis, such as finding out what the most used words are in the text.

3
00:00:09,200 --> 00:00:18,590
And we ended up with this list of tuples and we realized that articles are the words that are used the

4
00:00:18,590 --> 00:00:19,370
most.

5
00:00:21,230 --> 00:00:29,570
However, regular expressions the r e library doesn't allow us to distinguish if something is an article

6
00:00:29,570 --> 00:00:34,610
or it is something else such as a noun or a verb.

7
00:00:34,640 --> 00:00:38,560
And that is why we need an TC.

8
00:00:38,660 --> 00:00:41,090
The natural language processing for Python.

9
00:00:41,090 --> 00:00:46,910
So I will be opening a new Jupyter notebook and start the analysis.

10
00:00:47,910 --> 00:00:48,810
From scratch.

11
00:00:48,810 --> 00:00:54,300
I do recommend you do the same, although you can work on the previous notebook, whatever you prefer.

12
00:00:54,330 --> 00:00:55,560
Either way, it's fine.

13
00:00:56,640 --> 00:00:59,700
So in this notebook, I want to.

14
00:01:01,400 --> 00:01:04,700
Load the miracle in the end.

15
00:01:04,700 --> 00:01:06,500
This text file again.

16
00:01:07,040 --> 00:01:10,550
If you don't have this, find it attached in this lecture.

17
00:01:11,390 --> 00:01:13,490
So that will give us the book.

18
00:01:14,640 --> 00:01:17,540
As a String writes, that's the book.

19
00:01:17,560 --> 00:01:18,220
Now.

20
00:01:22,150 --> 00:01:24,640
So that was loading the book.

21
00:01:27,690 --> 00:01:35,580
Now let's find out the most used words non articles.

22
00:01:38,500 --> 00:01:42,670
Well to find the most use words which are not articles.

23
00:01:42,670 --> 00:01:49,240
First, we need to find the most used words, including articles as well.

24
00:01:49,240 --> 00:01:51,040
And we already did that.

25
00:01:53,540 --> 00:01:54,860
Somewhere in here.

26
00:01:56,000 --> 00:01:57,260
So using that.

27
00:02:02,090 --> 00:02:15,170
Codes there import or E, So we still need to use regular expressions to make a list of tuples like

28
00:02:15,170 --> 00:02:15,830
this.

29
00:02:15,830 --> 00:02:17,570
So that's called as well.

30
00:02:19,720 --> 00:02:21,820
So place that in another cell.

31
00:02:22,090 --> 00:02:24,400
And also we had this.

32
00:02:29,250 --> 00:02:31,980
So that's should return the findings here.

33
00:02:31,980 --> 00:02:35,490
I'm just printing out the first five words.

34
00:02:35,490 --> 00:02:41,220
So these are the words and then those words we convert them into a dictionary.

35
00:02:41,220 --> 00:02:49,830
So if I print out the dictionary here, we will get this the words along with the occurrences that they

36
00:02:49,830 --> 00:02:52,290
count in the texts in the book.

37
00:02:53,520 --> 00:02:56,730
So let me remove that and then we have this.

38
00:02:57,630 --> 00:03:01,980
So let's toll the list sorted.

39
00:03:03,410 --> 00:03:04,220
In here.

40
00:03:04,250 --> 00:03:06,680
The list is this.

41
00:03:07,550 --> 00:03:12,410
So with those three cells, we were able to construct a list of tuples.

42
00:03:12,560 --> 00:03:18,620
Each tuple contains the words and its number of occurrences in the text.

43
00:03:18,650 --> 00:03:19,430
No.

44
00:03:21,790 --> 00:03:26,560
Let me just display like the first five here.

45
00:03:27,220 --> 00:03:30,730
Now we need to use the analytic library.

46
00:03:33,880 --> 00:03:38,500
But first you need to install it to install it in a cell.

47
00:03:39,460 --> 00:03:45,890
You should do PIP 3.11 or 3.12, whatever Python version you have.

48
00:03:45,910 --> 00:03:47,650
Maybe you have 3.10.

49
00:03:48,680 --> 00:03:52,880
So you should know what version Jupiter Lab is using.

50
00:03:54,080 --> 00:04:02,300
If you don't know what Python version your Jupyter is using, you can do from platform import python

51
00:04:02,300 --> 00:04:03,770
underscore version.

52
00:04:04,550 --> 00:04:08,120
And then here do python underscore version.

53
00:04:08,120 --> 00:04:16,579
So call that function and that should return the python version you have then to again PIP 3.11 whatever

54
00:04:16,579 --> 00:04:24,560
version you had install nl TC that should install the analytical library in Jupyter.

55
00:04:27,540 --> 00:04:29,850
Then you should import and look.

56
00:04:32,180 --> 00:04:38,330
I went to from nlt k that corpus import.

57
00:04:39,820 --> 00:04:44,650
Stop words and then inside the variable.

58
00:04:47,360 --> 00:04:50,330
You want to stall the English stop words?

59
00:04:52,780 --> 00:04:54,310
For that you need to stop words.

60
00:04:54,310 --> 00:04:55,610
That words.

61
00:04:55,630 --> 00:05:00,910
Words is a method to download English stop words.

62
00:05:00,910 --> 00:05:02,230
So execute that.

63
00:05:04,790 --> 00:05:10,760
And then you want to check out what English stop words is.

64
00:05:11,030 --> 00:05:21,620
So these are things like articles and pronouns, so words which are used very often in English.

65
00:05:25,520 --> 00:05:27,770
And this is a list.

66
00:05:29,250 --> 00:05:32,910
So now we have that and we have this.

67
00:05:33,000 --> 00:05:42,480
Therefore, we can cross-check this against this and basically filter out the words of this lists so

68
00:05:42,480 --> 00:05:46,890
that this contains only the words which are not in this list.

69
00:05:47,700 --> 00:05:48,630
Let's do that.

70
00:05:49,980 --> 00:05:57,510
So we want to iterate four counts words in the list.

71
00:05:58,440 --> 00:06:01,680
So the list is this right, which contains.

72
00:06:02,490 --> 00:06:13,500
Those and counts will represent that number and words will represent those words that, that, that

73
00:06:14,160 --> 00:06:14,850
and so on.

74
00:06:16,650 --> 00:06:18,960
And when we say if words.

75
00:06:20,760 --> 00:06:23,400
Not in English.

76
00:06:23,400 --> 00:06:24,870
Stop words.

77
00:06:26,450 --> 00:06:28,910
And we want to create a new list.

78
00:06:28,910 --> 00:06:32,840
So let's create a new list here.

79
00:06:34,340 --> 00:06:38,900
Let's say filtered words is currently empty.

80
00:06:38,990 --> 00:06:47,450
But then as we iterate, we say filter two words that append and we append the words.

81
00:06:51,260 --> 00:06:59,210
Actually, we append a typical which contains the words and the count for each occurrence, which is

82
00:06:59,210 --> 00:07:01,040
not in this list.

83
00:07:01,040 --> 00:07:03,650
So if it's not in English, stop words.

84
00:07:04,950 --> 00:07:06,060
Execute that's.

85
00:07:06,240 --> 00:07:11,280
And then filtered words will be this.

86
00:07:13,700 --> 00:07:21,050
So now we see that the word out is actually the most used ones, then followed by us and said, and

87
00:07:21,050 --> 00:07:29,360
Roberto Roberto was the author's best friend, so we mentioned him a lot in the book.

88
00:07:29,660 --> 00:07:31,040
And then we have other words.

89
00:07:31,040 --> 00:07:40,100
Snow is very, very used here because in the book things are happening in a mountain covered, covered

90
00:07:40,100 --> 00:07:40,850
in snow.

91
00:07:41,120 --> 00:07:42,800
And there we have mountain as well.

92
00:07:43,460 --> 00:07:44,420
And so one.

93
00:07:47,810 --> 00:07:52,250
So that's how you filter out the words removing stop words.

94
00:07:52,250 --> 00:07:58,340
And so that's the first step to introducing natural language processing.

95
00:07:59,670 --> 00:08:09,020
So basically this library gives you access to words and their meanings, so to say.

96
00:08:09,030 --> 00:08:15,990
So in this case, we know that these are stop words, so words which are very commonly used like and,

97
00:08:15,990 --> 00:08:21,540
and that and etc. in the next videos we're going to be using.

98
00:08:22,840 --> 00:08:29,650
More examples of natural language processing so that you understand what it is exactly.

99
00:08:29,680 --> 00:08:30,790
I'll see you in the next video.

