1
00:00:01,930 --> 00:00:03,400
Welcome to this video.

2
00:00:03,430 --> 00:00:07,740
In this video, we will continue doing natural language processing.

3
00:00:07,750 --> 00:00:16,120
And particularly in this video, we will do sentiment analysis, and that is finding the mood of a piece

4
00:00:16,120 --> 00:00:18,520
of text if it's positive or negative.

5
00:00:19,090 --> 00:00:21,070
So let's see how we can do that.

6
00:00:21,190 --> 00:00:27,280
This is the code we wrote so for, and since this output is too long, I'll just print out just the

7
00:00:27,280 --> 00:00:35,020
first results of the most used words and then create a new cell here, Make it a modern cell.

8
00:00:35,020 --> 00:00:36,970
And this will be sentiment.

9
00:00:39,290 --> 00:00:40,310
Analysis.

10
00:00:42,100 --> 00:00:45,040
What is the most positive?

11
00:00:47,410 --> 00:00:49,060
The most negative.

12
00:00:50,350 --> 00:00:51,010
Chapter.

13
00:00:53,920 --> 00:00:56,320
To perform sentiment analysis.

14
00:00:56,320 --> 00:01:05,349
We need a sentiment intensity analyzer class, which comes from an ALT key dot sentiment.

15
00:01:06,400 --> 00:01:07,390
Import.

16
00:01:09,500 --> 00:01:12,500
Sentiment intensity analyzer.

17
00:01:15,300 --> 00:01:16,920
So I'm executing the sell.

18
00:01:16,920 --> 00:01:25,650
And in the next one, since we have just imported a class in here, we need to instantiate that class

19
00:01:25,650 --> 00:01:29,070
to create an object instance from that class.

20
00:01:31,520 --> 00:01:35,720
So sentiment intensity analyzer with parentheses.

21
00:01:36,640 --> 00:01:39,520
And that's who will create an analyzer.

22
00:01:41,590 --> 00:01:46,210
Object instance, which doesn't look any exciting for now.

23
00:01:46,210 --> 00:01:50,350
So you don't see any actual user friendly output.

24
00:01:51,460 --> 00:02:01,900
But this object instance now has methods you can check them with dear, and among them the one we need

25
00:02:02,650 --> 00:02:06,540
is a polarity score method.

26
00:02:06,550 --> 00:02:13,030
So that's polarity score.

27
00:02:14,200 --> 00:02:15,430
Polarity scores.

28
00:02:15,430 --> 00:02:16,060
Sorry.

29
00:02:17,140 --> 00:02:18,970
Yes, polarity scores.

30
00:02:20,970 --> 00:02:25,230
Polarity scores now expects a string.

31
00:02:25,500 --> 00:02:31,830
This could be one sentence, this could be one paragraphs, and this could also be an entire book.

32
00:02:31,950 --> 00:02:35,130
But it should be a string, for example, let's say.

33
00:02:37,380 --> 00:02:38,310
Hey, Luke.

34
00:02:38,490 --> 00:02:39,060
Ho!

35
00:02:40,110 --> 00:02:41,680
Through the trees.

36
00:02:41,700 --> 00:02:44,520
Ah, I love them.

37
00:02:48,680 --> 00:02:55,640
If you press control, enter now to execute this, you're going to get a dictionary as output.

38
00:02:59,730 --> 00:03:08,190
And these are the so called polarities, cause we have the negativity, the neutrality and the positivity

39
00:03:08,190 --> 00:03:10,830
and also a compound coefficient.

40
00:03:10,830 --> 00:03:12,040
So let's look at them.

41
00:03:12,060 --> 00:03:18,960
Negativity here is 0.0, obviously, because this is a very positive sentence.

42
00:03:19,320 --> 00:03:28,590
So we don't have a negativity there, but it can range from 0 to 1 and then we have neutrality.

43
00:03:28,710 --> 00:03:40,530
This is .46 for this can also range from 0 to 1 and then positivity is the highest score, so it's 0.53

44
00:03:40,710 --> 00:03:41,490
six.

45
00:03:42,270 --> 00:03:43,710
Then we have the compound.

46
00:03:43,710 --> 00:03:50,790
The compound might range from -1 to 1 and if it's above zero.

47
00:03:51,710 --> 00:03:58,970
As it is in this case, if it's above zero, that means that the sentence is more positive than negative.

48
00:03:58,970 --> 00:04:03,210
And if it's below zero, it means the sentence.

49
00:04:03,470 --> 00:04:08,060
The text is more negative than positive.

50
00:04:09,400 --> 00:04:15,710
So let's try to add something else here, such as I really love.

51
00:04:15,880 --> 00:04:19,029
So we're trying to make the sentence even more positive.

52
00:04:19,420 --> 00:04:24,820
If I execute that, we saw a slight increase in the positivity coefficient.

53
00:04:26,320 --> 00:04:33,280
But if we say, for example, how bad the trees are, I hate them and I really hate them, then we're

54
00:04:33,280 --> 00:04:40,960
going to see that the negativity has increased positivity zero and the compound coefficient is now negative.

55
00:04:42,340 --> 00:04:49,860
So in order to find out the moods of a text, you should use these polarity scores.

56
00:04:50,170 --> 00:05:00,910
Then perhaps you can write a conditional block, such as if Let's store this in schools, variable this

57
00:05:01,150 --> 00:05:03,460
dictionary and execute that.

58
00:05:04,380 --> 00:05:04,950
Right.

59
00:05:05,430 --> 00:05:07,800
And then you say if schools.

60
00:05:15,090 --> 00:05:19,140
Pause is greater than scores.

61
00:05:21,480 --> 00:05:24,270
Thank you.

62
00:05:24,270 --> 00:05:25,110
Print out.

63
00:05:26,560 --> 00:05:29,320
It is a positive text.

64
00:05:32,560 --> 00:05:34,420
Else print.

65
00:05:34,990 --> 00:05:37,870
It is a negative text.

66
00:05:37,870 --> 00:05:46,420
So basically, basically you can make use of those scores and make the program decide if a text is positive

67
00:05:46,420 --> 00:05:47,380
or negative.

68
00:05:48,970 --> 00:05:55,690
And of course, we can also try now this code on the book.

69
00:05:55,870 --> 00:06:02,860
So instead of that string, we can provide the entire variable press enter, and that should give you

70
00:06:02,860 --> 00:06:04,510
the dictionary with the.

71
00:06:05,430 --> 00:06:06,480
Scores.

72
00:06:07,750 --> 00:06:14,320
So since this book talks both about a tragedy but also a miracle, we do have a balance here.

73
00:06:14,320 --> 00:06:18,760
So the negativity is close to the positivity coefficient.

74
00:06:21,980 --> 00:06:25,760
Now let's try to analyze each chapter separately.

75
00:06:26,330 --> 00:06:28,550
So this here was just.

76
00:06:30,570 --> 00:06:31,680
An example.

77
00:06:33,180 --> 00:06:36,060
And now we need to analyze.

78
00:06:37,770 --> 00:06:38,920
Chapters.

79
00:06:40,200 --> 00:06:42,900
Sentiment Analysis.

80
00:06:46,560 --> 00:06:47,580
Right for that.

81
00:06:47,580 --> 00:06:54,390
We need the r e library and the pattern which should be r e that compile.

82
00:06:55,770 --> 00:06:58,560
And that was chapter.

83
00:07:00,710 --> 00:07:03,170
0 to 9 plus.

84
00:07:03,770 --> 00:07:10,340
So that will help us to split the book by this by chapters.

85
00:07:11,880 --> 00:07:17,340
We did cover regular expressions yesterday, so please go to yesterday's videos if you don't know what

86
00:07:17,340 --> 00:07:18,160
I'm doing here.

87
00:07:18,180 --> 00:07:22,770
Chapters will be are displayed.

88
00:07:22,770 --> 00:07:32,770
So the r e, the regular expression library does have a split method which splits a piece of text by

89
00:07:32,790 --> 00:07:37,290
pattern, so it will split the book by this pattern.

90
00:07:37,290 --> 00:07:40,470
Wherever this pattern is is met.

91
00:07:41,130 --> 00:07:45,480
That is where the split point will be on this book.

92
00:07:46,810 --> 00:07:53,530
So note chapters will be a list of chapters.

93
00:07:53,560 --> 00:08:01,990
However, the very first occurrence here is an empty string, so we need to get rid of that because

94
00:08:01,990 --> 00:08:09,040
you see, the length of chapters is not ten, but it is 11 because of that first string there.

95
00:08:10,090 --> 00:08:15,970
So we could say chapters is equal to chapters one.

96
00:08:15,980 --> 00:08:24,100
And so so we only extract the chapters starting from the first item from this very first chapter, excluding

97
00:08:24,100 --> 00:08:24,400
that.

98
00:08:24,400 --> 00:08:26,260
So execute that.

99
00:08:26,500 --> 00:08:32,620
Chapters now is the actual chapters right now.

100
00:08:34,340 --> 00:08:38,360
We can say from four chapter in chapters.

101
00:08:40,830 --> 00:08:44,159
Scores is equal to.

102
00:08:46,870 --> 00:08:49,420
That's methods.

103
00:08:50,950 --> 00:08:53,470
And in parentheses, we should provide the chapter.

104
00:08:53,470 --> 00:08:54,040
Right.

105
00:08:54,190 --> 00:08:55,810
Each chapter in here.

106
00:08:59,080 --> 00:09:01,060
And then print out the scores.

107
00:09:02,930 --> 00:09:04,460
So let's try this.

108
00:09:06,090 --> 00:09:07,050
And there we go.

109
00:09:07,470 --> 00:09:13,320
So we can look at the negativity here, so we can see, for example, that the very first chapter is

110
00:09:13,320 --> 00:09:17,310
the less negative of all, actually.

111
00:09:17,430 --> 00:09:25,140
And that is because the book talks about a team which was going to play a match somewhere else.

112
00:09:25,170 --> 00:09:30,360
And this was like full of harmony between the team players.

113
00:09:31,080 --> 00:09:33,930
They had great expectations of the game and so on.

114
00:09:36,000 --> 00:09:39,450
So there was no negativity in this chapter.

115
00:09:40,290 --> 00:09:48,030
But the positivity is also not the highest one because you see in this last chapter, the positivity

116
00:09:48,030 --> 00:09:49,170
was the highest.

117
00:09:49,470 --> 00:09:55,620
So this is the last chapter in the last chapter of the book, the ordeal ended.

118
00:09:56,190 --> 00:09:59,970
So there was a lot of positivity and so on.

119
00:10:01,810 --> 00:10:08,830
And I suppose that the middle chapters are the most negative one where the problems were happening in

120
00:10:08,830 --> 00:10:09,880
the book and so on.

121
00:10:11,110 --> 00:10:11,530
All right.

122
00:10:11,530 --> 00:10:14,620
So that's what you get.

123
00:10:14,890 --> 00:10:16,870
Of course, you can take this even further.

124
00:10:16,870 --> 00:10:25,180
You can perhaps write some effects conditionals or you can even write full number.

125
00:10:26,140 --> 00:10:37,660
So the chapter in our chapter in enumerate chapters, and then you can print out the number along the

126
00:10:37,660 --> 00:10:38,470
scores.

127
00:10:39,670 --> 00:10:40,600
So there you go.

128
00:10:40,630 --> 00:10:43,210
Chapter number zero, chapter number one and so on.

129
00:10:43,210 --> 00:10:45,550
Or you can say number plus one.

130
00:10:47,170 --> 00:10:53,770
And you get the correct numbers of the chapters along the Sentiment Analysis dictionary.

131
00:10:54,820 --> 00:10:59,780
So with that, I thank you for following this and I'll talk to you in the next videos.

