Python实现文件查询关键字功能的示例详解

 更新时间:2026年02月12日 09:52:17   作者:豆本-豆豆奶  
这篇文章主要为大家详细介绍了Python实现文件查询关键字功能,文中的示例详解详细,具有一定的借鉴价值,感兴趣的小伙伴可以跟随小编一起学习一下

咱们可以想象一个这样的场景,你这边有大量的文件,现在我需要查找一个 序列号:xxxxxx,我想知道这个 序列号:xxxxxx 在哪个文件中。

在没有使用代码脚本的情况下,你可能需要一个文件一个文件打开,然后按 CTRL+F 来进行搜索查询。

那既然我们会使用 python,何不自己写一个呢?本文将实现这样一个工具,且源码全在文章中,只需要复制粘贴即可使用。

思路

主要思路就是通过打开文件夹,获取文件,一个个遍历查找关键字,流程图如下:

流程图

怎么样,思路非常简单,所以其实实现也不难。

本文将支持少部分文件类型,更多类型需要读者自己实现:

  • txt
  • docx
  • csv
  • xlsx
  • pptx

读取txt

安装库

pip install chardet

代码

import chardet


def detect_encoding(file_path):
    raw_data = None
    with open(file_path, 'rb') as f:
        for line in f:
            raw_data = line
            break

        if raw_data is None:
            raw_data = f.read()
    result = chardet.detect(raw_data)
    return result['encoding']


def read_txt(file_path, keywords=''):
    is_in = False
    encoding = detect_encoding(file_path)
    with open(file_path, 'r', encoding=encoding) as f:
        for line in f:
            if line.find(keywords) != -1:
                is_in = True
                break

    return is_in

我们使用了 chardet 库来判断 txt 的编码,以应对不同编码的读取方式。

读取docx

安装库

pip install python-docx

代码

from docx import Document


def read_docx(file_path, keywords=''):
    doc = Document(file_path)
    is_in = False

    for para in doc.paragraphs:
        if para.text.find(keywords) != -1:
            is_in = True
            break

    return is_in

读取csv

代码

import csv


def read_csv(file_path, keywords=''):
    is_in = False

    encoding = detect_encoding(file_path)
    with open(file_path, mode='r', encoding=encoding) as f:
        reader = csv.reader(f)

        for row in reader:
            row_text = ''.join([str(v) for v in row])
            if row_text.find(keywords) != -1:
                is_in = True
                break

    return is_in

读取xlsx

安装库

pip install openpyxl

代码

from openpyxl import load_workbook


def read_xlsx(file_path, keywords=''):
    wb = load_workbook(file_path)
    sheet_names = wb.sheetnames

    is_in = False
    for sheet_name in sheet_names:
        sheet = wb[sheet_name]
        for row in sheet.iter_rows(values_only=True):
            row_text = ''.join([str(v) for v in row])
            if row_text.find(keywords) != -1:
                is_in = True
                break

    wb.close()

    return is_in

读取pptx

安装库

pip install python-pptx 

代码

from pptx import Presentation


def read_ppt(ppt_file, keywords=''):
    prs = Presentation(ppt_file)
    is_in = False
    for slide in prs.slides:
        for shape in slide.shapes:
            if shape.has_text_frame:
                text_frame = shape.text_frame
                for paragraph in text_frame.paragraphs:
                    for run in paragraph.runs:
                        if run.text.find(keywords) != -1:
                            is_in = True
                            break

    return is_in

文件夹递归

为了防止文件夹嵌套导致的问题,我们还有一个文件夹递归的操作。

代码

from pathlib import Path


def list_files_recursive(directory):
    file_paths = []

    for path in Path(directory).rglob('*'):
        if path.is_file():
            file_paths.append(str(path))

    return file_paths

完整代码

# -*- coding: utf-8 -*-
from pptx import Presentation
import chardet
from docx import Document
import csv
from openpyxl import load_workbook
from pathlib import Path


def detect_encoding(file_path):
    raw_data = None
    with open(file_path, 'rb') as f:
        for line in f:
            raw_data = line
            break

        if raw_data is None:
            raw_data = f.read()
    result = chardet.detect(raw_data)
    return result['encoding']


def read_txt(file_path, keywords=''):
    is_in = False
    encoding = detect_encoding(file_path)
    with open(file_path, 'r', encoding=encoding) as f:
        for line in f:
            if line.find(keywords) != -1:
                is_in = True
                break

    return is_in


def read_docx(file_path, keywords=''):
    doc = Document(file_path)
    is_in = False

    for para in doc.paragraphs:
        if para.text.find(keywords) != -1:
            is_in = True
            break

    return is_in


def read_csv(file_path, keywords=''):
    is_in = False

    encoding = detect_encoding(file_path)
    with open(file_path, mode='r', encoding=encoding) as f:
        reader = csv.reader(f)

        for row in reader:
            row_text = ''.join([str(v) for v in row])
            if row_text.find(keywords) != -1:
                is_in = True
                break

    return is_in


def read_xlsx(file_path, keywords=''):
    wb = load_workbook(file_path)
    sheet_names = wb.sheetnames

    is_in = False
    for sheet_name in sheet_names:
        sheet = wb[sheet_name]
        for row in sheet.iter_rows(values_only=True):
            row_text = ''.join([str(v) for v in row])
            if row_text.find(keywords) != -1:
                is_in = True
                break

    wb.close()

    return is_in


def read_ppt(ppt_file, keywords=''):
    prs = Presentation(ppt_file)
    is_in = False
    for slide in prs.slides:
        for shape in slide.shapes:
            if shape.has_text_frame:
                text_frame = shape.text_frame
                for paragraph in text_frame.paragraphs:
                    for run in paragraph.runs:
                        if run.text.find(keywords) != -1:
                            is_in = True
                            break

    return is_in


def list_files_recursive(directory):
    file_paths = []

    for path in Path(directory).rglob('*'):
        if path.is_file():
            file_paths.append(str(path))

    return file_paths


if __name__ == '__main__':
    keywords = '测试关键字'
    file_paths = list_files_recursive(r'测试文件夹')
    for file_path in file_paths:
        if file_path.endswith('.txt'):
            is_in = read_txt(file_path, keywords)
        elif file_path.endswith('.docx'):
            is_in = read_docx(file_path, keywords)
        elif file_path.endswith('.csv'):
            is_in = read_csv(file_path, keywords)
        elif file_path.endswith('.xlsx'):
            is_in = read_xlsx(file_path, keywords)
        elif file_path.endswith('.pptx'):
            is_in = read_ppt(file_path, keywords)

        if is_in:
            print(file_path)

结尾

现在你可以十分方便地使用代码查找出各种文件中是否存在关键字了

以上就是Python实现文件查询关键字功能的示例详解的详细内容,更多关于Python查询文件关键字的资料请关注脚本之家其它相关文章!

相关文章

  • python斐波那契数列的计算方法

    python斐波那契数列的计算方法

    这篇文章主要为大家详细介绍了python斐波那契数列的计算方法,具有一定的参考价值,感兴趣的小伙伴们可以参考一下
    2018-09-09
  • PyTorch加载数据集梯度下降优化

    PyTorch加载数据集梯度下降优化

    这篇文章主要介绍了PyTorch加载数据集梯度下降优化,使用DataLoader方法,并继承DataSet抽象类,可实现对数据集进行mini_batch梯度下降优化,需要的小伙伴可以参考一下
    2022-03-03
  • Python logging日志模块的核心用法与实操技巧

    Python logging日志模块的核心用法与实操技巧

    logging 是 Python 标准库中的一个模块,它提供了灵活的日志记录功能,通过 logging,开发者可以方便地将日志信息输出到控制台、文件、网络等多种目标,本文给大家介绍了Python logging日志模块的核心用法与实操技巧,需要的朋友可以参考下
    2026-03-03
  • PyQt5每天必学之日历控件QCalendarWidget

    PyQt5每天必学之日历控件QCalendarWidget

    这篇文章主要为大家详细介绍了PyQt5每天必学之日历控件QCalendarWidget,具有一定的参考价值,感兴趣的小伙伴们可以参考一下
    2018-04-04
  • Python Process创建进程的2种方法详解

    Python Process创建进程的2种方法详解

    这篇文章主要介绍了Python Process创建进程的2种方法详解,文中通过示例代码介绍的非常详细,对大家的学习或者工作具有一定的参考学习价值,需要的朋友们下面随着小编来一起学习学习吧
    2021-01-01
  • Kmeans均值聚类算法原理以及Python如何实现

    Kmeans均值聚类算法原理以及Python如何实现

    这个算法中文名为k均值聚类算法,首先我们在二维的特殊条件下讨论其实现的过程,方便大家理解。
    2020-09-09
  • python爬虫开发之Request模块从安装到详细使用方法与实例全解

    python爬虫开发之Request模块从安装到详细使用方法与实例全解

    这篇文章主要介绍了python爬虫开发之Request模块从安装到详细使用方法与实例全解,需要的朋友可以参考下
    2020-03-03
  • 深入浅析Python中的迭代器

    深入浅析Python中的迭代器

    迭代器是实现了迭代器协议的类对象,迭代器协议规定了迭代器类必需定义__next()__方法。这篇文章主要介绍了Python中的迭代器,需要的朋友可以参考下
    2019-06-06
  • Python有序查找算法之二分法实例分析

    Python有序查找算法之二分法实例分析

    这篇文章主要介绍了Python有序查找算法之二分法,结合实例形式分析了Python二分查找算法的原理与相关实现技巧,需要的朋友可以参考下
    2017-12-12
  • Python实现HTTP网络请求功能的入门指南

    Python实现HTTP网络请求功能的入门指南

    HTTP是互联网上应用最广泛的通信协议,简单来说,HTTP 网络请求就是客户端(如你的 Python 程序)向服务器发送消息,并等待服务器返回响应的过程,本文给大家介绍了在Python中实现HTTP网络请求功能的入门指南,需要的朋友可以参考下
    2026-05-05

最新评论