removing duplicates rows of a dataframe python

Question

I need to remove duplicate rows from a dataset. Basically, I should perform

proc sort data=mydata noduprecs dupout=mydata_dup;run;

I need to remove duplicates as well as save those duplicate rows in a separate dataframe . How can I do that?

Answer 1

Assuming your dataset is a pandas dataframe.

To remove the duplicated rows:

data = data.drop_duplicates()

To select all the duplicated rows:

dup = data.ix[data.duplicated(), :]

Hope it helps.

Answer 2

A few examples from Pandas docs :

> df = pd.DataFrame({

    'brand': ['Yum Yum', 'Yum Yum', 'Indomie', 'Indomie', 'Indomie'],

    'style': ['cup', 'cup', 'cup', 'pack', 'pack'],

    'rating': [4, 4, 3.5, 15, 5]

})

> df
    brand style  rating
0  Yum Yum   cup     4.0
1  Yum Yum   cup     4.0
2  Indomie   cup     3.5
3  Indomie  pack    15.0
4  Indomie  pack     5.0

By default, it removes duplicate rows based on all columns.

> df.drop_duplicates()
    brand style  rating
0  Yum Yum   cup     4.0
2  Indomie   cup     3.5
3  Indomie  pack    15.0
4  Indomie  pack     5.0

To remove duplicates on specific column(s), use subset.

> df.drop_duplicates(subset=['brand'])
    brand style  rating
0  Yum Yum   cup     4.0
2  Indomie   cup     3.5

To remove duplicates and keep last occurrences, use keep.

> df.drop_duplicates(subset=['brand', 'style'], keep='last')
    brand style  rating
1  Yum Yum   cup     4.0
2  Indomie   cup     3.5
4  Indomie  pack     5.0

removing duplicates rows of a dataframe python

Question

2 answers

solution1
0 ACCPTED 2017-07-06 18:02:44

solution2
0 2021-02-26 22:37:30

removing duplicates rows of a dataframe python

Question

2 answers

solution1 0 ACCPTED 2017-07-06 18:02:44

solution2 0 2021-02-26 22:37:30

solution1
0 ACCPTED 2017-07-06 18:02:44

solution2
0 2021-02-26 22:37:30