xlsx
anthropics/skills
只要試算表檔案是主要輸入或輸出,即可隨時使用此技能。 這意味著任何使用者希望執行以下操作的任務:開啟、讀取、編輯或修正現有的 .xlsx、.xlsm、.xltx、.csv 或 .tsv 檔案(例如:新增欄位、運算公式、格式設定、製作圖表、清理雜亂資料); 從頭開始建立新的試算表,或從其他資料來源建立試算表;或是將表格檔案格式相互轉換。特別是在使用者透過名稱或路徑引用試算表檔案時(即使是隨口一提,例如「我下載資料夾裡的那個 xlsx 檔案」),且希望對該檔案進行某項操作或從中產生輸出時,應觸發此技能。 此外,當需清理或重整混亂的表格資料檔案(例如格式錯誤的行、位置錯誤的標題列、垃圾資料)並將其轉為正確的試算表時,亦會觸發此條件。最終產出物必須為試算表檔案。 若主要交付成果為 Word 文件、HTML 報告、獨立的 Python 腳本、資料庫處理流程,或 Google 試算表 API 整合,即使涉及表格資料,亦請勿觸發此任務。
...展開全部XLSX 建立、編輯與分析
openpyxl、pandas和markitdown已預先安裝 — 請勿先執行pip install;請直接撰寫腳本並導入模組。僅當導入失敗(或缺少markitdown指令)時,才需執行pip install 安裝缺失的套件。
以下腳本路徑均以本技能的目錄為基準。
所有輸出的要求
- 除非使用者另有指定,否則全文均須使用專業字型(Arial、Times New Roman)。
- 公式錯誤為零。若
recalc.py報告有 errors_found,絕不發布。若您認為錯誤是您接手前就存在的,請提供證明:以data_only=True載入原始檔案,並檢視該儲存格。您引入的錯誤與繼承而來的錯誤外觀完全相同。 - 請使用公式,切勿硬編碼結果。應寫作
sheet['B10'] = '=SUM(B2:B9)',而非直接使用 Python 計算出的總和。當輸入值變更時,試算表必須重新計算。 - 嚴格遵循使用者的規格說明。包括精確的標籤名稱、精確的欄位標題,以及他們所指定的公式。即使再優雅,若重新設計後計算的內容與原規格不符,即屬失敗。
- 請在讀者可見之處(例如儲存格註解,或表格末端的相鄰儲存格)記錄每項假設與硬編碼的數值。 若有實際來源,請註明(
來源:公司 10-K 報告,2024 財政年度,第 45 頁,營收說明,[SEC EDGAR 網址]);若數字來自使用者,請明確說明。 - 若您建立的工作簿是供他人填寫,則需附上簡短說明,標明哪些儲存格可供編輯,並提供一行包含實際數值的範例,以展示預期格式。切勿在受託編輯的檔案中新增此類範例行。
- 編輯現有檔案時:務必完全遵循其既定規範。這些規範優先於本指南中的所有指示。請先找出指定的輸入儲存格——通常以獨特的字體顏色、底色或陰影標示——僅在這些位置輸入資料,並保持所有現有公式不變。
重新計算(只要檔案包含公式,此步驟即為強制要求)
openpyxl 會將公式寫入為字串,且不包含快取值。在您重新計算之前,任何
讀取快取值的程式(例如pandas、
load_workbook(data_only=True) 以及大多數預覽工具)讀取公式儲存格時,都會顯示為None。
python 腳本/recalc.py 輸出。xlsx [timeout_seconds] # 預設 30
LibreOffice 會計算每個公式,檔案會就地重新寫入,您將獲得 JSON:
status(success|errors_found)、total_formulas、total_errors,以及一個
error_summary,其中每種錯誤類型最多列出 100 個儲存格(locations_truncated表示被
省略的數量 — 請以total_errors 為準,而非清單長度)。 修正其標示的錯誤並再次執行。
若JSON 中的status 鍵被errors鍵取代,表示未進行任何重新計算,且
僅該情況會返回非零值 —errors_found會返回 0,因此切勿將正常退出視為工作簿
無誤。
綠色重新計算結果僅證明公式能成功評估,並不代表公式正確。若範圍偏移一個單元格 或引用了錯誤的列,雖會產生乾淨且無錯誤的檔案,但其中的數字卻是錯誤的。 在建立整個網格之前,請先撰寫 2 至 3 個公式,並確認它們能取得您預期的數值。
若工作簿連結至其他檔案,當您使用 openpyxl 重新儲存後
再重新計算,這些連結將會遺失。 此類公式的格式為='[1]Returns Analysis'!$B$2—— 其中的[1]是
工作簿外部參照清單中的索引,用以指定磁碟上的獨立檔案,而非工作表。
由於該檔案在此處極少存在,因此該儲存格的快取值是唯一保存其
資料的來源。openpyxl 會在儲存時移除該值;LibreOffice 隨後必須實際解析該參照,
但因解析失敗而寫入#NAME? 並刪除所有連結。recalc.py在此狀態下將拒絕執行
——請在覆寫儲存前,先將這些儲存格的值從原始檔案複製出來(--force參數會覆寫此設定,
並接受資料遺失的後果)。
選擇能通過驗證的公式
LibreOffice 實作的功能比 Excel 少,而任何無法評估的函數都會變成
嵌入您交付檔案中的#NAME?字面值。
- 請優先選用 Excel 2007 時期的函數—
SUMIFS、INDEX、MATCH、IFERROR、SUMPRODUCT— 這些函數無需前綴。 - 有六種 2007 年後的函數雖然可用,但必須加上
_xlfn.前綴,因為 openpyxl 會將您的公式原樣寫入 XML,而 Excel 儲存 2007 年後的函數名稱時會加上前綴(其使用者介面會隱藏該前綴):_xlfn.TEXTJOIN、_xlfn.CONCAT、_xlfn.IFS、_xlfn.SWITCH、_xlfn.MAXIFS、_xlfn.MINIFS。若直接書寫,每個都會產生#NAME?錯誤。 - 切勿使用
XLOOKUP、XMATCH、SORT、FILTER、UNIQUE或SEQUENCE。執行環境中的 LibreOffice 無論使用何種前綴,都無法評估這些函數。 較新的版本雖然能評估這些函數,但它們屬於溢出陣列函數,而由 openpyxl 撰寫的檔案沒有溢出元資料,因此只有範圍的左上角儲存格會獲得數值——且recalc.py針對這份被截斷的結果會回報total_errors: 0。 請使用INDEX/MATCH進行查詢,並在寫入儲存格之前,先在 Python 中執行排序、篩選及去重。 - LibreOffice 無法解析的公式在寫回時會轉為小寫——這是除了
#NAME?之外的另一項快速辨識指標。
openpyxl 的注意事項
- 讀取模型需要進行兩次載入。
設定 data_only=True會產生已移除公式的快取值;預設設定則會產生不含數值的公式字串。單次處理無法同時取得這兩者。 - 若在儲存前設定 `
data_only=True`,將造成資料破壞。該工作簿中將不剩任何公式,因此儲存時會將每個公式永久替換為字面值。 - 若對 openpyxl 剛寫入的檔案
設定 data_only=True,所有位置都會返回None—— 請先執行recalc.py。(結果為""的公式讀取時也會返回None。) - 合併儲存格:僅寫入左上角的錨點。範圍內的其他所有儲存格皆為
MergedCell,其.value屬性為唯讀。 - 除非在
load_workbook時傳入keep_vba=True,否則.xlsm檔案會遺失其巨集。 - 若工作表名稱包含空格,在跨工作表參照時必須加上引號:
='Assumptions Inputs'!$B$5。若未加引號,將評估為#VALUE!。
財務模型
除非使用者另有指示,或現有檔案已有其他處理方式。
顏色:藍色文字 (0,0,255) 用於硬編碼輸入值和情境調整因子 · 黑色用於公式 ·
連結至其他工作表的單元格為綠色 (0,128,0) · 連結至其他檔案的單元格為紅色 (255,0,0) ·
關鍵假設及需由使用者填入的儲存格則以黃色填充 (255,255,0)。
數字:貨幣格式$#,##0,單位名稱顯示於標題中(收入 ($mm)) · 零
顯示為-,百分比亦同 ($#,##0;($#,##0);-) · 負數以括號標示 ·
百分比0.0 %,儲存為分數(0.15顯示為15.0%;儲存15則顯示為
1500.0%) · 估值倍數0.0x· 年份以文字顯示(「2024」,絕不顯示為2,024)。
結構:每個假設值皆置於獨立且標有標籤的儲存格中,並由使用該假設值的公式進行引用
(=B5*(1+$B$6),絕不使用=B5*1.05) · 所有預測期間的公式保持一致,因為
行中單一修訂的儲存格是最常見的隱性錯誤 · 對可能為零的分母進行保護。
依賴項
openpyxl、pandas、markitdown(pip,預先安裝 — 僅在匯入失敗或缺少指令時才需安裝) · LibreOffice(soffice,透過scripts/office/soffice.py 自動配置為沙盒環境)
XLSX creation, editing, and analysis
openpyxl,pandas, andmarkitdownare preinstalled — do not runpip installfirst; write the script and import directly. Only if an import fails (or themarkitdowncommand is missing):pip installthe missing package.
Script paths below are relative to this skill's directory.
Requirements for every output
- Professional font (Arial, Times New Roman) throughout, unless the user says otherwise.
- Zero formula errors. Never ship while
recalc.pyreportserrors_found. If you think an error predates you, prove it: load the original withdata_only=Trueand look at that cell. An error you introduced looks exactly like one you inherited. - Use formulas, never hardcoded results. Write
sheet['B10'] = '=SUM(B2:B9)', not the Python-computed total. The sheet must recalculate when its inputs change. - Follow the user's spec literally. Exact tab names, exact column headers, and the formula they spelled out. A redesign that computes something else fails, however elegant.
- Document every assumption and hardcoded number where the reader will see it — a cell comment, or an adjacent cell at a table's end. Cite a real source when one exists (
Source: Company 10-K, FY2024, Page 45, Revenue Note, [SEC EDGAR URL]); when the number came from the user, say so plainly. - A workbook you create for someone to fill in needs a short legend naming which cells to edit, and one example row of realistic values showing the expected format. Never add such a row to a file you were asked to edit.
- Editing an existing file: match its conventions exactly. They override every guideline here. Find its designated input cells first — a distinct font color, fill, or shading marks them — write only there, and leave every existing formula untouched.
Recalculate (mandatory whenever the file contains formulas)
openpyxl writes formulas as strings with no cached values. Until you recalculate, every
formula cell reads back as None to anything reading cached values — pandas,
load_workbook(data_only=True), and most previewers.
python scripts/recalc.py output.xlsx [timeout_seconds] # default 30
LibreOffice computes every formula, the file is rewritten in place, and you get JSON:
status (success | errors_found), total_formulas, total_errors, and an
error_summary naming up to 100 cells per error type (locations_truncated says how many it
withheld — trust total_errors, not the length of the list). Fix what it names and run it
again. JSON with an error key instead of a status means nothing was recalculated, and
only that case exits non-zero — errors_found exits 0, so never treat a clean exit as a clean
workbook.
A green recalc proves your formulas evaluate, not that they are right. An off-by-one range or a reference to the wrong row yields a clean, error-free file with wrong numbers. Write 2–3 formulas first and check they pull the values you expect, before building out a grid.
A workbook that links to another file loses those links if you re-save it with openpyxl and
then recalculate. Such a formula reads ='[1]Returns Analysis'!$B$2 — the [1] is an index
into the workbook's external-reference list, naming a separate file on disk, not a sheet.
That file is rarely present here, so the cell's cached value is the only thing holding its
data. openpyxl strips that value on save; LibreOffice then has to resolve the reference for
real, fails, writes #NAME?, and deletes every link. recalc.py refuses to run in that state
— copy those cells' values out of the original before you save over them (--force overrides,
and accepts the loss).
Choosing formulas that survive verification
LibreOffice implements fewer functions than Excel, and one it cannot evaluate becomes a
literal #NAME? baked into the file you deliver.
- Prefer Excel-2007-era functions —
SUMIFS,INDEX,MATCH,IFERROR,SUMPRODUCT— which need no prefix. - Six post-2007 functions work, but only with an
_xlfn.prefix, because openpyxl writes your formula into the XML verbatim and Excel stores post-2007 names prefixed (its UI hides the prefix):_xlfn.TEXTJOIN,_xlfn.CONCAT,_xlfn.IFS,_xlfn.SWITCH,_xlfn.MAXIFS,_xlfn.MINIFS. Written bare, each yields#NAME?. - Never use
XLOOKUP,XMATCH,SORT,FILTER,UNIQUE, orSEQUENCE. The runtime's LibreOffice cannot evaluate them under any prefix. Newer builds do evaluate them, but they are spilling array functions and an openpyxl-written file has no spill metadata, so only the top-left cell of the range gets a value — andrecalc.pyreportstotal_errors: 0on the truncated result. UseINDEX/MATCHfor lookups, and sort, filter, and de-duplicate in Python before writing the cells. - A formula LibreOffice could not parse is written back lowercased — a quick tell beside a
#NAME?.
openpyxl gotchas
- Reading a model takes two loads.
data_only=Trueyields cached values with the formulas gone; the default yields formula strings with no values. One pass cannot give you both. data_only=Trueis destructive if you save. That workbook has no formulas left, so saving replaces every one with a literal — permanently.data_only=Trueon a file openpyxl just wrote returnsNoneeverywhere — runrecalc.pyfirst. (A formula whose result is""also reads back asNone.)- Merged cells: write the top-left anchor only. Every other cell in the range is a
MergedCellwhose.valueis read-only. .xlsmloses its macros unless you passkeep_vba=Truetoload_workbook.- A sheet name containing a space must be quoted in a cross-sheet reference:
='Assumptions Inputs'!$B$5. Unquoted, it evaluates to#VALUE!.
Financial models
Unless the user says otherwise, or the existing file already does something else.
Color: blue text (0,0,255) for hardcoded inputs and scenario levers · black for formulas ·
green (0,128,0) for links to another sheet · red (255,0,0) for links to another file ·
yellow fill (255,255,0) for key assumptions and cells the user should fill in.
Numbers: currency $#,##0, with the unit named in the header (Revenue ($mm)) · zeros
render as -, including in percentages ($#,##0;($#,##0);-) · negatives in parentheses ·
percentages 0.0%, stored as fractions (0.15 renders 15.0%; storing 15 renders
1500.0%) · valuation multiples 0.0x · years as text ("2024", never 2,024).
Structure: every assumption in its own labeled cell, referenced by the formulas that use it
(=B5*(1+$B$6), never =B5*1.05) · formulas consistent across every projection period, since a
lone edited cell mid-row is the commonest silent error · guard denominators that can be zero.
Dependencies
openpyxl, pandas, markitdown (pip, preinstalled — install only if an import fails or the command is missing) · LibreOffice (soffice, auto-configured for sandboxed environments via scripts/office/soffice.py)





首頁
