Tag: Web scraping

  • 3 ขั้นตอนในการทำ Web Scraping ด้วย BeautifulSoup ใน Python

    3 ขั้นตอนในการทำ Web Scraping ด้วย BeautifulSoup ใน Python

    Web scraping เป็นการดึงข้อมูลจากเว็บไซต์ด้วย code หรือระบบอัตโนมัติ

    เราสามารถใช้ประโยชน์จาก web scraping ได้หลากหลาย เช่น:

    • เปรียบเทียบราคา
    • วิเคราะห์ brand
    • เพิ่มข้อมูลให้กับ AI

    ในบทความนี้ ผมจะพาทุกคนไปดูวิธีทำ web scraping ด้วย BeautifulSoup ใน Python ซึ่งเป็น package ยอดนิยมสำหรับเริ่มต้นทำ web scraping กัน

    ถ้าพร้อมแล้ว ไปเริ่มกันเลย



    🪜 ภาพรวม

    การทำ web scraping ด้วย BeautifulSoup มีอยู่ 3 ขั้นตอน:

    1. Get the web page
    2. Create a Soup object
    3. Extract the content

    เราไปดูรายละเอียดแต่ละขั้น ผ่านตัวอย่างการดึงข้อมูลจาก Book to Scrape ซึ่งเป็นเว็บไซต์สำหรับฝึก web scraping โดยเฉพาะกัน


    1️⃣ Step 1. Get the Web Page

    ในขั้นแรก เราจะโหลดหน้าเว็บที่ต้องการ

    ในตัวอย่าง เราจะโหลดหน้าเว็บหนังสือ Sapiens: A Brief History of Humankind กัน:

    BeautifulSoup ไม่มี function สำหรับโหลดหน้าเว็บในตัว ดังนั้น เราจะต้องใช้ package อื่นอย่าง requests เข้ามาช่วย:

    # Load the package
    import requests
    # Set the URL
    url = "<https://books.toscrape.com/catalogue/sapiens-a-brief-history-of-humankind_996/index.html>"
    # Get the response
    response = requests.get(url)
    # Encode the response
    response.encoding = "utf-8"

    Note: เรากำหนด response.encoding = "utf-8” เพื่อให้อ่านอักขระพิเศษ (เช่น £) ได้

    เพื่อเช็กว่าเราโหลดหน้าเว็บสำเร็จไหม เราสามารถดูจาก status code ได้:

    # Print the status code
    print(response.status_code)

    ผลลัพธ์:

    200

    ในตัวอย่าง status code 200 หมายถึง โหลดหน้าเว็บสำเร็จ


    2️⃣ Step 2. Create a Soup Object

    ในขั้นที่ 2 เราจะเปลี่ยนหน้าเว็บที่โหลดได้ให้เป็น Soup object ที่เราจะสามารถดึงข้อมูลได้ง่าย:

    # Import the package
    import bs4
    # Create Soup
    soup = bs4.BeautifulSoup(response.text, "html.parser")

    เราสามารถดูตัวอย่าง Soup object ได้แบบนี้:

    # Print the first 1,000 characters
    print(soup.prettify()[:1000])

    ผลลัพธ์:

    <!DOCTYPE html>
    <!--[if lt IE 7]> <html lang="en-us" class="no-js lt-ie9 lt-ie8 lt-ie7"> <![endif]-->
    <!--[if IE 7]> <html lang="en-us" class="no-js lt-ie9 lt-ie8"> <![endif]-->
    <!--[if IE 8]> <html lang="en-us" class="no-js lt-ie9"> <![endif]-->
    <!--[if gt IE 8]><!-->
    <html class="no-js" lang="en-us">
    <!--<![endif]-->
    <head>
    <title>
    Sapiens: A Brief History of Humankind | Books to Scrape - Sandbox
    </title>
    <meta content="text/html; charset=utf-8" http-equiv="content-type"/>
    <meta content="24th Jun 2016 09:29" name="created"/>
    <meta content="
    From a renowned historian comes a groundbreaking narrative of humanity’s creation and evolution—a #1 international bestseller—that explores the ways in which biology and history have defined us and enhanced our understanding of what it means to be “human.”One hundred thousand years ago, at least six different species of humans inhabited Earth. Yet today there is only one—h From a renowned historian comes a g

    สังเกตว่า Soup ประกอบด้วยเนื้อหาและ HTML tag ของหน้าเว็บ


    3️⃣ Step 3. Extract the Content

    ในขั้นสุดท้าย เมื่อได้ Soup object แล้ว เราสามารถดึงข้อมูลที่ต้องการได้ง่าย ๆ ด้วย method อย่าง:

    • .find() สำหรับค้นหา HTML tag ที่ต้องการ
    • .get_text() สำหรับดึงเนื้อหาที่เป็น text
    • .get() สำหรับดึง attribute จาก HTML tag

    ตัวอย่างเช่น ดึงชื่อ:

    # Get the book title
    title = soup.find("h1").get_text()
    # Print the title
    print(title)

    ผลลัพธ์:

    Sapiens: A Brief History of Humankind

    ดึงราคา:

    # Get the price
    price = soup.find("p", class_="price_color").get_text()
    # Print the price
    print(price)

    ผลลัพธ์:

    £54.23

    หรือดึงรูปภาพหนังสือ:

    # Get the img tag
    image_relative_url = soup.find("img").get("src")
    # Set base URL
    base_url = "<https://books.toscrape.com/>"
    # Concatenate the image URL
    image_full_url = base_url + image_relative_url.replace("../", "")
    # Print the image
    print(image_full_url)

    ผลลัพธ์:

    <https://books.toscrape.com/media/cache/ce/5f/ce5f052c65cc963cf4422be096e915c9.jpg>

    เมื่อเปิด URL แล้ว เราจะได้ภาพหนังสือแบบนี้:


    💪 บทสรุป

    วิธีใช้ BeautifulSoup สำหรับ web scraping มีอยู่ 3 ขั้นตอน:

    ขั้นที่ 1. Get the web page:

    # Import the package
    import requests
    # Set the URL
    url = "<https://books.toscrape.com/catalogue/sapiens-a-brief-history-of-humankind_996/index.html>"
    # Get the response
    response = requests.get(url)

    ขั้นที่ 2. Create a Soup object:

    # Import the package
    import bs4
    # Create Soup
    soup = bs4.BeautifulSoup(response.text, "html.parser")

    ขั้นที่ 3. Extract the content:

    # Get the book title
    title = soup.find("h1").get_text()
    # Print the title
    print(title)

    🫵 หลังจบบทความนี้

    หลังอ่านบทความนี้แล้ว อย่าลืมไปลองทำ web scraping กันนะครับ:


    📃 อ้างอิง