簡體   English   中英

在Pandas DataFrame中創建名稱列

[英]Create Names column in Pandas DataFrame

我使用的是Python包names以生成QA測試一些名字。

names包中包含函數names.get_first_name(gender) ,它允許字符串male或female作為參數。 目前我有以下DataFrame:

    Marital Gender
0   Single  Female
1   Married Male
2   Married Male
3   Single  Male
4   Married Female

我嘗試過以下方法:

df.loc[df.Gender == 'Male', 'FirstName'] = names.get_first_name(gender = 'male')
df.loc[df.Gender == 'Female', 'FirstName'] = names.get_first_name(gender = 'female')

但我收到的只是兩個名字:

    Marital Gender  FirstName
0   Single  Female  Kathleen
1   Married Male    David
2   Married Male    David
3   Single  Male    David
4   Married Female  Kathleen

有沒有辦法分別為每一行調用此函數,因此並非所有男性/女性都具有相同的確切名稱?

你需要申請

 df['Firstname']=df['Gender'].str.lower().apply(names.get_first_name)

您可以使用列表理解:

df['Firstname']= [names.get_first_name(gender) for gender in df['Gender'].str.lower()] 

聽到是一個黑客,按性別(連同他們的概率)讀取所有名稱,然后隨機抽樣。

import names

def get_names(gender):
    if not isinstance(gender, (str, unicode)) or gender.lower() not in ('male', 'female'):
        raise ValueError('Invalid gender')

    with open(names.FILES['first:{}'.format(gender.lower())], 'rb') as fin:
        first_names = []
        probs = []
        for line in fin:
            first_name, prob, dummy, dummy = line.strip().split()
            first_names.append(first_name)
            probs.append(float(prob) / 100)
    return pd.DataFrame({'first_name': first_names, 'probability': probs})

def get_random_first_names(n, first_names_by_gender):
    first_names = (
        first_names_by_gender
        .sample(n, replace=True, weights='probability')
        .loc[:, 'first_name']
        .tolist()
    )
    return first_names

first_names = {gender: get_names(gender) for gender in ('Male', 'Female')}

>>> get_random_first_names(3, first_names['Male'])
['RICHARD', 'EDWARD', 'HOMER']

>>> get_random_first_names(4, first_names['Female'])
['JANICE', 'CAROLINE', 'DOROTHY', 'DIANE']

如果使用map速度很重要

list(map(names.get_first_name,df.Gender))
Out[51]: ['Harriett', 'Parker', 'Alfred', 'Debbie', 'Stanley']
#df['FN']=list(map(names.get_first_name,df.Gender))

暫無
暫無

聲明:本站的技術帖子網頁,遵循CC BY-SA 4.0協議,如果您需要轉載,請注明本站網址或者原文地址。任何問題請咨詢:yoyou2525@163.com.

 
粵ICP備18138465號  © 2020-2024 STACKOOM.COM