3 ขั้นตอนในการทำ Web Scraping ด้วย BeautifulSoup ใน Python

BeautifulSoup (web scrapping package in Python)

Web scraping เป็นการดึงข้อมูลจากเว็บไซต์ด้วย code หรือระบบอัตโนมัติ

เราสามารถใช้ประโยชน์จาก web scraping ได้หลากหลาย เช่น:

  • เปรียบเทียบราคา
  • วิเคราะห์ brand
  • เพิ่มข้อมูลให้กับ AI

ในบทความนี้ ผมจะพาทุกคนไปดูวิธีทำ web scraping ด้วย BeautifulSoup ใน Python ซึ่งเป็น package ยอดนิยมสำหรับเริ่มต้นทำ web scraping กัน

ถ้าพร้อมแล้ว ไปเริ่มกันเลย



🪜 ภาพรวม

การทำ web scraping ด้วย BeautifulSoup มีอยู่ 3 ขั้นตอน:

  1. Get the web page
  2. Create a Soup object
  3. Extract the content

เราไปดูรายละเอียดแต่ละขั้น ผ่านตัวอย่างการดึงข้อมูลจาก Book to Scrape ซึ่งเป็นเว็บไซต์สำหรับฝึก web scraping โดยเฉพาะกัน


1️⃣ Step 1. Get the Web Page

ในขั้นแรก เราจะโหลดหน้าเว็บที่ต้องการ

ในตัวอย่าง เราจะโหลดหน้าเว็บหนังสือ Sapiens: A Brief History of Humankind กัน:

BeautifulSoup ไม่มี function สำหรับโหลดหน้าเว็บในตัว ดังนั้น เราจะต้องใช้ package อื่นอย่าง requests เข้ามาช่วย:

# Load the package
import requests
# Set the URL
url = "<https://books.toscrape.com/catalogue/sapiens-a-brief-history-of-humankind_996/index.html>"
# Get the response
response = requests.get(url)
# Encode the response
response.encoding = "utf-8"

Note: เรากำหนด response.encoding = "utf-8” เพื่อให้อ่านอักขระพิเศษ (เช่น £) ได้

เพื่อเช็กว่าเราโหลดหน้าเว็บสำเร็จไหม เราสามารถดูจาก status code ได้:

# Print the status code
print(response.status_code)

ผลลัพธ์:

200

ในตัวอย่าง status code 200 หมายถึง โหลดหน้าเว็บสำเร็จ


2️⃣ Step 2. Create a Soup Object

ในขั้นที่ 2 เราจะเปลี่ยนหน้าเว็บที่โหลดได้ให้เป็น Soup object ที่เราจะสามารถดึงข้อมูลได้ง่าย:

# Import the package
import bs4
# Create Soup
soup = bs4.BeautifulSoup(response.text, "html.parser")

เราสามารถดูตัวอย่าง Soup object ได้แบบนี้:

# Print the first 1,000 characters
print(soup.prettify()[:1000])

ผลลัพธ์:

<!DOCTYPE html>
<!--[if lt IE 7]> <html lang="en-us" class="no-js lt-ie9 lt-ie8 lt-ie7"> <![endif]-->
<!--[if IE 7]> <html lang="en-us" class="no-js lt-ie9 lt-ie8"> <![endif]-->
<!--[if IE 8]> <html lang="en-us" class="no-js lt-ie9"> <![endif]-->
<!--[if gt IE 8]><!-->
<html class="no-js" lang="en-us">
<!--<![endif]-->
<head>
<title>
Sapiens: A Brief History of Humankind | Books to Scrape - Sandbox
</title>
<meta content="text/html; charset=utf-8" http-equiv="content-type"/>
<meta content="24th Jun 2016 09:29" name="created"/>
<meta content="
From a renowned historian comes a groundbreaking narrative of humanity’s creation and evolution—a #1 international bestseller—that explores the ways in which biology and history have defined us and enhanced our understanding of what it means to be “human.”One hundred thousand years ago, at least six different species of humans inhabited Earth. Yet today there is only one—h From a renowned historian comes a g

สังเกตว่า Soup ประกอบด้วยเนื้อหาและ HTML tag ของหน้าเว็บ


3️⃣ Step 3. Extract the Content

ในขั้นสุดท้าย เมื่อได้ Soup object แล้ว เราสามารถดึงข้อมูลที่ต้องการได้ง่าย ๆ ด้วย method อย่าง:

  • .find() สำหรับค้นหา HTML tag ที่ต้องการ
  • .get_text() สำหรับดึงเนื้อหาที่เป็น text
  • .get() สำหรับดึง attribute จาก HTML tag

ตัวอย่างเช่น ดึงชื่อ:

# Get the book title
title = soup.find("h1").get_text()
# Print the title
print(title)

ผลลัพธ์:

Sapiens: A Brief History of Humankind

ดึงราคา:

# Get the price
price = soup.find("p", class_="price_color").get_text()
# Print the price
print(price)

ผลลัพธ์:

£54.23

หรือดึงรูปภาพหนังสือ:

# Get the img tag
image_relative_url = soup.find("img").get("src")
# Set base URL
base_url = "<https://books.toscrape.com/>"
# Concatenate the image URL
image_full_url = base_url + image_relative_url.replace("../", "")
# Print the image
print(image_full_url)

ผลลัพธ์:

<https://books.toscrape.com/media/cache/ce/5f/ce5f052c65cc963cf4422be096e915c9.jpg>

เมื่อเปิด URL แล้ว เราจะได้ภาพหนังสือแบบนี้:


💪 บทสรุป

วิธีใช้ BeautifulSoup สำหรับ web scraping มีอยู่ 3 ขั้นตอน:

ขั้นที่ 1. Get the web page:

# Import the package
import requests
# Set the URL
url = "<https://books.toscrape.com/catalogue/sapiens-a-brief-history-of-humankind_996/index.html>"
# Get the response
response = requests.get(url)

ขั้นที่ 2. Create a Soup object:

# Import the package
import bs4
# Create Soup
soup = bs4.BeautifulSoup(response.text, "html.parser")

ขั้นที่ 3. Extract the content:

# Get the book title
title = soup.find("h1").get_text()
# Print the title
print(title)

🫵 หลังจบบทความนี้

หลังอ่านบทความนี้แล้ว อย่าลืมไปลองทำ web scraping กันนะครับ:


📃 อ้างอิง

Comments

Leave a Reply

Discover more from Shi no Shigoto | AI, Data, & Psychology

Subscribe now to keep reading and get access to the full archive.

Continue reading