Tag: 抜き出し

  • Extracting Specific Values from a DataFrame [Python]

    Extracting Specific Values from a DataFrame [Python]

    Extracting a single specific value from a pandas DataFrame in Python is surprisingly tricky.
    DataFrames are designed to be handled as DataFrames, so pulling out a row, slicing, or appending data is very easy.

    But when you try to extract a value, for some reason it does not come out as a value—it stays in DataFrame form, and various other problems come up.

    Here I describe how to extract the value in a DataFrame that satisfies a particular condition.
    There are probably better ways to do this, but this is the best I can manage at my current level.

    First, suppose we have the DataFrame below and want to extract the value 20 outlined in red.

    The catch is that with a simple DataFrame like the one above you can see the index and pull the value out in one step, but with a large dataset it is very hard to check the index number and use it.

    With a large dataset, it is more efficient to extract data using a key column (customer number, ID, etc.) as a clue.
    So let’s try extracting the value in the situation below.

    In other words, suppose we want the value in column B for the row where column A is Aichi.
    You cannot do this in one step; the value is extracted in two stages.

    First, create a DataFrame like the one above.

    import pandas as pd
    df = pd.DataFrame({ 'A' : ["Tokyo", "Aichi", "Osaka"],
                        'B' : [10, 20, 30],
                        'C' : [100, 200, 300]})
    df

    Next, extract the row where column A is Aichi.

    a = df[df["A"] == "Aichi"]
    
    print(a)

    This lets us pull out the row containing Aichi.
    Finally, we extract the value in column B of this row.

    b = a.at[a.index[0], "B"]
    
    print(b)

    Doing this, you can extract the value.
    It is roundabout, but this is the method I currently use.

    If you know of another approach, I would be glad to hear about it in the comments.

    import pandas as pd
    df = pd.DataFrame({ 'A' : ["Tokyo", "Aichi", "Osaka"],
                        'B' : [10, 20, 30],
                        'C' : [100, 200, 300]})
    
    a = df[df["A"] == "Aichi"]
    b = a.at[a.index[0], "B"]
    print(b)

    Here is a summary of the code above.