@@ -7,8 +7,9 @@ This model is used for Sentiment Analysis, which is based on BERTurk for Turkish
...
@@ -7,8 +7,9 @@ This model is used for Sentiment Analysis, which is based on BERTurk for Turkish
# Dataset
# Dataset
We used product and movie dataset provided by the study [2] . This dataset includes
The dataset is taken from the studies [2] and [3] and merged.
movie and product reviews. The products are book, DVD, electronics, and kitchen.
* The study [2] gathered movie and product reviews. The products are book, DVD, electronics, and kitchen.
The movie dataset is taken from a cinema Web page (www.beyazperde.com) with
The movie dataset is taken from a cinema Web page (www.beyazperde.com) with
5331 positive and 5331 negative sentences. Reviews in the Web page are marked in
5331 positive and 5331 negative sentences. Reviews in the Web page are marked in
scale from 0 to 5 by the users who made the reviews. The study considered a review
scale from 0 to 5 by the users who made the reviews. The study considered a review
...
@@ -18,16 +19,70 @@ Web page. They constructed benchmark dataset consisting of reviews regarding som
...
@@ -18,16 +19,70 @@ Web page. They constructed benchmark dataset consisting of reviews regarding som
products (book, DVD, etc.). Likewise, reviews are marked in the range from 1 to 5,
products (book, DVD, etc.). Likewise, reviews are marked in the range from 1 to 5,
and majority class of reviews are 5. Each category has 700 positive and 700 negative
and majority class of reviews are 5. Each category has 700 positive and 700 negative
reviews in which average rating of negative reviews is 2.27 and of positive reviews
reviews in which average rating of negative reviews is 2.27 and of positive reviews
is 4.5.
is 4.5. This dataset is also used the study [1]
* The study[3] collected tweet dataset. They proposed a new approach for automatically classifying the sentiment of microblog messages. The proposed approach is based on utilizing robust feature representation and fusion.
The dataset is used by following papers
*Merged Dataset*
| *size* | *data* |
|--------|----|
| 8000 |dev.tsv|
| 8262 |test.tsv|
| 32000 |train.tsv|
| *48290* |*total*|
The dataset is used by following papers
* 1 Yildirim, Savaş. (2020). Comparing Deep Neural Networks to Traditional Models for Sentiment Analysis in Turkish Language. 10.1007/978-981-15-1216-2_12.
* 1 Yildirim, Savaş. (2020). Comparing Deep Neural Networks to Traditional Models for Sentiment Analysis in Turkish Language. 10.1007/978-981-15-1216-2_12.
* 2 Demirtas, Erkin and Mykola Pechenizkiy. 2013. Cross-lingual polarity detection with machine translation. In Proceedings of the Second International Workshop on Issues of Sentiment
* 2 Demirtas, Erkin and Mykola Pechenizkiy. 2013. Cross-lingual polarity detection with machine translation. In Proceedings of the Second International Workshop on Issues of Sentiment
Discovery and Opinion Mining (WISDOM ’13)
Discovery and Opinion Mining (WISDOM ’13)
* Hayran, A., Sert, M. (2017), "Sentiment Analysis on Microblog Data based on Word Embedding and Fusion Techniques", IEEE 25th Signal Processing and Communications Applications Conference (SIU 2017), Belek, Turkey