Web scraping เป็นการดึงข้อมูลจากเว็บไซต์ด้วย code หรือระบบอัตโนมัติ
เราสามารถใช้ประโยชน์จาก web scraping ได้หลากหลาย เช่น:
- เปรียบเทียบราคา
- วิเคราะห์ brand
- เพิ่มข้อมูลให้กับ AI
ในบทความนี้ ผมจะพาทุกคนไปดูวิธีทำ web scraping ด้วย BeautifulSoup ใน Python ซึ่งเป็น package ยอดนิยมสำหรับเริ่มต้นทำ web scraping กัน
ถ้าพร้อมแล้ว ไปเริ่มกันเลย
🪜 ภาพรวม
การทำ web scraping ด้วย BeautifulSoup มีอยู่ 3 ขั้นตอน:
- Get the web page
- Create a
Soupobject - Extract the content
เราไปดูรายละเอียดแต่ละขั้น ผ่านตัวอย่างการดึงข้อมูลจาก Book to Scrape ซึ่งเป็นเว็บไซต์สำหรับฝึก web scraping โดยเฉพาะกัน
1️⃣ Step 1. Get the Web Page
ในขั้นแรก เราจะโหลดหน้าเว็บที่ต้องการ
ในตัวอย่าง เราจะโหลดหน้าเว็บหนังสือ Sapiens: A Brief History of Humankind กัน:

BeautifulSoup ไม่มี function สำหรับโหลดหน้าเว็บในตัว ดังนั้น เราจะต้องใช้ package อื่นอย่าง requests เข้ามาช่วย:
# Load the packageimport requests# Set the URLurl = "<https://books.toscrape.com/catalogue/sapiens-a-brief-history-of-humankind_996/index.html>"# Get the responseresponse = requests.get(url)# Encode the responseresponse.encoding = "utf-8"
Note: เรากำหนด response.encoding = "utf-8” เพื่อให้อ่านอักขระพิเศษ (เช่น £) ได้
เพื่อเช็กว่าเราโหลดหน้าเว็บสำเร็จไหม เราสามารถดูจาก status code ได้:
# Print the status codeprint(response.status_code)
ผลลัพธ์:
200
ในตัวอย่าง status code 200 หมายถึง โหลดหน้าเว็บสำเร็จ
2️⃣ Step 2. Create a Soup Object
ในขั้นที่ 2 เราจะเปลี่ยนหน้าเว็บที่โหลดได้ให้เป็น Soup object ที่เราจะสามารถดึงข้อมูลได้ง่าย:
# Import the packageimport bs4# Create Soupsoup = bs4.BeautifulSoup(response.text, "html.parser")
เราสามารถดูตัวอย่าง Soup object ได้แบบนี้:
# Print the first 1,000 charactersprint(soup.prettify()[:1000])
ผลลัพธ์:
<!DOCTYPE html><!--[if lt IE 7]> <html lang="en-us" class="no-js lt-ie9 lt-ie8 lt-ie7"> <![endif]--><!--[if IE 7]> <html lang="en-us" class="no-js lt-ie9 lt-ie8"> <![endif]--><!--[if IE 8]> <html lang="en-us" class="no-js lt-ie9"> <![endif]--><!--[if gt IE 8]><!--><html class="no-js" lang="en-us"> <!--<![endif]--> <head> <title> Sapiens: A Brief History of Humankind | Books to Scrape - Sandbox </title> <meta content="text/html; charset=utf-8" http-equiv="content-type"/> <meta content="24th Jun 2016 09:29" name="created"/> <meta content=" From a renowned historian comes a groundbreaking narrative of humanity’s creation and evolution—a #1 international bestseller—that explores the ways in which biology and history have defined us and enhanced our understanding of what it means to be “human.”One hundred thousand years ago, at least six different species of humans inhabited Earth. Yet today there is only one—h From a renowned historian comes a g
สังเกตว่า Soup ประกอบด้วยเนื้อหาและ HTML tag ของหน้าเว็บ
3️⃣ Step 3. Extract the Content
ในขั้นสุดท้าย เมื่อได้ Soup object แล้ว เราสามารถดึงข้อมูลที่ต้องการได้ง่าย ๆ ด้วย method อย่าง:
.find()สำหรับค้นหา HTML tag ที่ต้องการ.get_text()สำหรับดึงเนื้อหาที่เป็น text.get()สำหรับดึง attribute จาก HTML tag
ตัวอย่างเช่น ดึงชื่อ:
# Get the book titletitle = soup.find("h1").get_text()# Print the titleprint(title)
ผลลัพธ์:
Sapiens: A Brief History of Humankind
ดึงราคา:
# Get the priceprice = soup.find("p", class_="price_color").get_text()# Print the priceprint(price)
ผลลัพธ์:
£54.23
หรือดึงรูปภาพหนังสือ:
# Get the img tagimage_relative_url = soup.find("img").get("src")# Set base URLbase_url = "<https://books.toscrape.com/>"# Concatenate the image URLimage_full_url = base_url + image_relative_url.replace("../", "")# Print the imageprint(image_full_url)
ผลลัพธ์:
<https://books.toscrape.com/media/cache/ce/5f/ce5f052c65cc963cf4422be096e915c9.jpg>
เมื่อเปิด URL แล้ว เราจะได้ภาพหนังสือแบบนี้:

💪 บทสรุป
วิธีใช้ BeautifulSoup สำหรับ web scraping มีอยู่ 3 ขั้นตอน:
ขั้นที่ 1. Get the web page:
# Import the packageimport requests# Set the URLurl = "<https://books.toscrape.com/catalogue/sapiens-a-brief-history-of-humankind_996/index.html>"# Get the responseresponse = requests.get(url)
ขั้นที่ 2. Create a Soup object:
# Import the packageimport bs4# Create Soupsoup = bs4.BeautifulSoup(response.text, "html.parser")
ขั้นที่ 3. Extract the content:
# Get the book titletitle = soup.find("h1").get_text()# Print the titleprint(title)
🫵 หลังจบบทความนี้
หลังอ่านบทความนี้แล้ว อย่าลืมไปลองทำ web scraping กันนะครับ:

Leave a Reply