Commits · 4b372f3f7e4455fa96253089b6a592adc664f3e2 · wangsen / MinerU

"magic_pdf/vscode:/vscode.git/clone" did not exist on "56fad23d67a654a2b033e36638ee35095e5c8a51"

09 Sep, 2024 1 commit

feat(ocr): pass language parameter for custom model init · 4b372f3f

myhloli authored Sep 09, 2024

Pass the `lang` parameter to `custom_model_init` in `doc_analyze` to support language-specific OCR configurations. This enhancement allows the use of language information to improve OCR accuracy when processing PDFs.

4b372f3f

30 Aug, 2024 1 commit

feat(cli&analyze&pipeline): add start_page and end_page args for pagination (#507) · 0f91fcf6

Xiaomeng Zhao authored Aug 30, 2024

* feat(cli&analyze&pipeline): add start_page and end_page args for paginationAdd start_page_id and end_page_id arguments to various components of the PDF parsing
pipeline to support pagination functionality. This feature allows users to specify the
range of pages to be processed, enhancing the efficiency and flexibility of the system.

* feat(cli&analyze&pipeline): add start_page and end_page args for paginationAdd start_page_id and end_page_id arguments to various components of the PDF parsing
pipeline to support pagination functionality. This feature allows users to specify the
range of pages to be processed, enhancing the efficiency and flexibility of the system.

* feat(cli&analyze&pipeline): add start_page and end_page args for paginationAdd start_page_id and end_page_id arguments to various components of the PDF parsing
pipeline to support pagination functionality. This feature allows users to specify the
range of pages to be processed, enhancing the efficiency and flexibility of the system.

0f91fcf6

28 Aug, 2024 1 commit
- fix: remove the default value of output option in tools/cli.py and tools/cli_dev.py (#494) · f0a8886c
  icecraft authored Aug 28, 2024
```
Co-authored-by: icecraft <xurui1@pjlab.org.cn>
```
  f0a8886c
20 Aug, 2024 2 commits

fix(ocr_mkcontent): revise table caption output (#397) · dd19f59e

Xiaomeng Zhao authored Aug 20, 2024



* fix(ocr_mkcontent): revise table caption output

- Ensuring that
  table captions are properly included in the output.
- Remove the redundant `table_caption` variable。

* Update cla.yml

* Update bug_report.yml

* feat(cli): add debug option for detailed error handling

Enable users to invoke the CLI command with a new debug flag to get detailed debugging information.

* fix(pdf-extract-kit): adjust crop_paste parameters for better accuracyThe crop_paste_x and crop_paste_y values in the pdf_extract_kit.py have been modified
to improve the accuracy and consistency of OCR processing. The new values are set to 25
to ensure more precise image cropping and pasting which leads to better OCR recognition
results.

* Update README_zh-CN.md (#404)

correct FAQ url

* Update README_zh-CN.md (#404) (#409) (#410)

correct FAQ url
Co-authored-by: sfk <18810651050@163.com>

* Update FAQ_zh_cn.md

add new issue

* Update FAQ_en_us.md

* Update README_Windows_CUDA_Acceleration_zh_CN.md

* Update README_zh-CN.md

* @Thepathakarpit has signed the CLA in opendatalab/MinerU#418

* fix(pdf-extract-kit): increase crop_paste margin for OCR processingDouble the crop_paste margin from25 to 50 to ensure better OCR accuracy and
handling of border cases. This change will help in improving the overall quality of
OCR'ed text by providing more context around the detected text areas.

* fix(common): deep copy model list before drawing model bbox

Use a deep copy of the original model list in `drow_model_bbox` to avoid potential
modifications to the source data. This ensures the integrity of the original models
is maintained while generating the model bounding boxes visualization.

---------
Co-authored-by: sfk <18810651050@163.com>
Co-authored-by: drunkpig <60862764+drunkpig@users.noreply.github.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>

dd19f59e

feat: rename the file generated by command line tools (#401) · c9a51491

icecraft authored Aug 20, 2024



* feat: rename the file generated by command line tools

* feat: add pdf filename as prefix to {span,layout,model}.pdf

---------
Co-authored-by: icecraft <tmortred@gmail.com>
Co-authored-by: icecraft <xurui1@pjlab.org.cn>

c9a51491

01 Aug, 2024 2 commits

feat: remove dummpy code, magic_pdf/cli, magic_pdf/train_utils (#291) · e155d322

icecraft authored Aug 01, 2024



* feat: remove dummpy code, magic_pdf/cli, magic_pdf/train_utils

* feat: expose version in command line

---------
Co-authored-by: shenguanlin <shenguanlin@pjlab.org.cn>

e155d322

Feat/impl cli (#264) · 40e0827e

icecraft authored Aug 01, 2024



* feat: refractor cli command

* feat: add docs to describe the output files of cli

* feat: resove review comments

* feat: updat docs about middle.json

---------
Co-authored-by: shenguanlin <shenguanlin@pjlab.org.cn>

40e0827e