代码之家  ›  专栏  ›  技术社区  ›  Gary Paluk

Z80 ASM BNF结构…我走对了吗?

  •  10
  • Gary Paluk  · 技术社区  · 17 年前

    我正在尝试学习bnf并尝试组装一些z80 asm代码。既然我对这两个领域都不熟悉,我的问题是,我是否走上了正确的道路?我正在尝试将z80 asm的格式写为ebnf,这样我就可以从那里找到从源代码创建机器代码的位置。目前我有以下几点:

    Assignment = Identifier, ":" ;
    
    Instruction = Opcode, [ Operand ], [ Operand ] ;
    
    Operand = Identifier | Something* ;
    
    Something* = "(" , Identifier, ")" ;
    
    Identifier = Alpha, { Numeric | Alpha } ;
    
    Opcode = Alpha, Alpha ;
    
    Int = [ "-" ], Numeric, { Numeric } ;
    
    Alpha = "A" | "B" | "C" | "D" | "E" | "F" | 
            "G" | "H" | "I" | "J" | "K" | "L" | 
            "M" | "N" | "O" | "P" | "Q" | "R" | 
            "S" | "T" | "U" | "V" | "W" | "X" | 
            "Y" | "Z" ;
    
    Numeric = "0" | "1" | "2" | "3"| "4" | 
              "5" | "6" | "7" | "8" | "9" ;
    

    如果我出错了,任何方向性的反馈都是很好的。

    3 回复  |  直到 13 年前
        1
  •  17
  •   Ira Baxter    17 年前

    传统的汇编程序通常是在汇编程序中手工编码的,并使用特殊的解析技术来处理汇编源代码行以生成实际的汇编程序代码。 当汇编程序语法很简单(例如,总是操作码寄存器、操作数)时,这就足够有效了。

    现代机器的指令集杂乱不堪,有许多指令变化和操作数,它们可以用复杂的语法来表示,允许多个索引寄存器参与操作数表达式。允许使用具有固定和可重定位常量的复杂的装配时间表达式以及各种类型的加法运算符,这会使情况复杂化。允许条件编译、宏、结构化数据声明等的复杂的汇编程序都增加了对语法的新要求。用特殊方法处理所有这些语法是非常困难的,这也是解析器生成器被发明的原因。

    使用BNF和解析器生成器是构建现代汇编程序的非常合理的方法,即使对于Z80这样的遗留处理器也是如此。我已经为摩托罗拉的8位机器(如6800/6809)构建了这样的汇编程序,并准备为现代x86做同样的工作。我认为你正朝着正确的方向前进。

    **********编辑********** OP要求提供lexer和parser定义示例。 我在这里都提供了。

    这些是6809装配机的实际规格摘录。 完整的定义是这里样本大小的2-3倍。

    为了减少空间,我删去了很多黑暗角落的复杂性。 这就是这些定义的意义所在。 一个人可能会对异常的复杂感到沮丧; 关键是,有了这样的定义,你就可以 描述 这个 语言的形状,而不是按程序编码。 如果你 以一种特别的方式对所有这些代码进行编码,这将远远超过 可维护性较低。

    了解这些定义也会有所帮助 与高端程序分析系统一起使用, 将词法/分析工具作为子系统,称为 The DMS Software Reengineering Toolkit . DMS将自动从
    语法规则在解析器规范中,这使它成为 更容易构建解析工具。最后, 解析器规范包含所谓的“预打印器” 声明,允许DMS从AST重新创建源文本。 (语法的真正目的是让我们建立代表汇编程序的ASTS 指令,然后把它们吐出来给真正的汇编程序!)

    注意一点:词法和语法规则是如何表述的(元语法!) 在不同的lexer/parser生成器系统之间有所不同。这个 基于DMS规范的语法也不例外。DMS相对复杂 它本身的语法规则,在这里可用的空间里,这真的不太实际。你必须接受其他系统使用类似符号的想法,因为 ebnf用于规则和lexems的正则表达式变体。

    考虑到操作人员的兴趣,他可以实现类似的lexer/parsers 使用任何lexer/parser生成器工具,例如flex/yacc, Javacc,Antlr,…

    **********莱克斯***

    -- M6809.lex: Lexical Description for M6809
    -- Copyright (C) 1989,1999-2002 Ira D. Baxter
    
    %%
    #mainmode Label
    
    #macro digit "[0-9]"
    #macro hexadecimaldigit "<digit>|[a-fA-F]"
    
    #macro comment_body_character "[\u0009 \u0020-\u007E]" -- does not include NEWLINE
    
    #macro blank "[\u0000 \ \u0009]"
    
    #macro hblanks "<blank>+"
    
    #macro newline "\u000d \u000a? \u000c? | \u000a \u000c?" -- form feed allowed only after newline
    
    #macro bare_semicolon_comment "\; <comment_body_character>* "
    
    #macro bare_asterisk_comment "\* <comment_body_character>* "
    
    ...[snip]
    
    #macro hexadecimal_digit "<digit> | [a-fA-F]"
    
    #macro binary_digit "[01]"
    
    #macro squoted_character "\' [\u0021-\u007E]"
    
    #macro string_character "[\u0009 \u0020-\u007E]"
    
    %%Label -- (First mode) processes left hand side of line: labels, opcodes, etc.
    
    #skip "(<blank>*<newline>)+"
    #skip "(<blank>*<newline>)*<blank>+"
      << (GotoOpcodeField ?) >>
    
    #precomment "<comment_line><newline>"
    
    #preskip "(<blank>*<newline>)+"
    #preskip "(<blank>*<newline>)*<blank>+"
      << (GotoOpcodeField ?) >>
    
    -- Note that an apparant register name is accepted as a label in this mode
    #token LABEL [STRING] "<identifier>"
      <<  (local (;; (= [TokenScan natural] 1) ; process all string characters
             (= [TokenLength natural] ?:TokenCharacterCount)=
             (= [TokenString (reference TokenBodyT)] (. ?:TokenCharacters))
             (= [Result (reference string)] (. ?:Lexeme:Literal:String:Value))
             [ThisCharacterCode natural]
             (define Ordinala #61)
             (define Ordinalf #66)
             (define OrdinalA #41)
             (define OrdinalF #46)
         );;
         (;; (= (@ Result) `') ; start with empty string
         (while (<= TokenScan TokenLength)
          (;;   (= ThisCharacterCode (coerce natural TokenString:TokenScan))  
            (+= TokenScan) ; bump past character
            (ifthen (>= ThisCharacterCode Ordinala)
               (-= ThisCharacterCode #20) ; fold to upper case
            )ifthen
            (= (@ Result) (append (@ Result) (coerce character ThisCharacterCode)))=
    
            );;
         )while
         );;
      )local
      (= ?:Lexeme:Literal:String:Format (LiteralFormat:MakeCompactStringLiteralFormat 0))  ; nothing interesting in string
      (GotoLabelList ?)
      >>
    
    %%OpcodeField
    
    #skip "<hblanks>"
      << (GotoEOLComment ?) >>
    #ifnotoken
      << (GotoEOLComment ?) >>
    
    -- Opcode field tokens
    #token 'ABA'       "[aA][bB][aA]"
       << (GotoEOLComment ?) >>
    #token 'ABX'       "[aA][bB][xX]"
       << (GotoEOLComment ?) >>
    #token 'ADC'       "[aA][dD][cC]"
       << (GotoABregister ?) >>
    #token 'ADCA'      "[aA][dD][cC][aA]"
       << (GotoOperand ?) >>
    #token 'ADCB'      "[aA][dD][cC][bB]"
       << (GotoOperand ?) >>
    #token 'ADCD'      "[aA][dD][cC][dD]"
       << (GotoOperand ?) >>
    #token 'ADD'       "[aA][dD][dD]"
       << (GotoABregister ?) >>
    #token 'ADDA'      "[aA][dD][dD][aA]"
       << (GotoOperand ?) >>
    #token 'ADDB'      "[aA][dD][dD][bB]"
       << (GotoOperand ?) >>
    #token 'ADDD'      "[aA][dD][dD][dD]"
       << (GotoOperand ?) >>
    #token 'AND'       "[aA][nN][dD]"
       << (GotoABregister ?) >>
    #token 'ANDA'      "[aA][nN][dD][aA]"
       << (GotoOperand ?) >>
    #token 'ANDB'      "[aA][nN][dD][bB]"
       << (GotoOperand ?) >>
    #token 'ANDCC'     "[aA][nN][dD][cC][cC]"
       << (GotoRegister ?) >>
    ...[long list of opcodes snipped]
    
    #token IDENTIFIER [STRING] "<identifier>"
      <<  (local (;; (= [TokenScan natural] 1) ; process all string characters
             (= [TokenLength natural] ?:TokenCharacterCount)=
             (= [TokenString (reference TokenBodyT)] (. ?:TokenCharacters))
             (= [Result (reference string)] (. ?:Lexeme:Literal:String:Value))
             [ThisCharacterCode natural]
             (define Ordinala #61)
             (define Ordinalf #66)
             (define OrdinalA #41)
             (define OrdinalF #46)
         );;
         (;; (= (@ Result) `') ; start with empty string
         (while (<= TokenScan TokenLength)
          (;;   (= ThisCharacterCode (coerce natural TokenString:TokenScan))  
            (+= TokenScan) ; bump past character
            (ifthen (>= ThisCharacterCode Ordinala)
               (-= ThisCharacterCode #20) ; fold to upper case
            )ifthen
            (= (@ Result) (append (@ Result) (coerce character ThisCharacterCode)))=
    
            );;
         )while
         );;
      )local
      (= ?:Lexeme:Literal:String:Format (LiteralFormat:MakeCompactStringLiteralFormat 0))  ; nothing interesting in string
      (GotoOperandField ?)
      >>
    
    #token '#'   "\#" -- special constant introduction (FDB)
       << (GotoDataField ?) >>
    
    #token NUMBER [NATURAL] "<decimal_number>"
      << (local [format LiteralFormat:NaturalLiteralFormat]
        (;; (= ?:Lexeme:Literal:Natural:Value (ConvertDecimalTokenStringToNatural (. format) ? 0 0))
        (= ?:Lexeme:Literal:Natural:Format (LiteralFormat:MakeCompactNaturalLiteralFormat format))
        );;
     )local
     (GotoOperandField ?)
      >>
    
    #token NUMBER [NATURAL] "\$ <hexadecimal_digit>+"
      << (local [format LiteralFormat:NaturalLiteralFormat]
        (;; (= ?:Lexeme:Literal:Natural:Value (ConvertHexadecimalTokenStringToNatural (. format) ? 1 0))
        (= ?:Lexeme:Literal:Natural:Format (LiteralFormat:MakeCompactNaturalLiteralFormat format))
        );;
     )local
     (GotoOperandField ?)
      >>
    
    #token NUMBER [NATURAL] "\% <binary_digit>+"
      << (local [format LiteralFormat:NaturalLiteralFormat]
        (;; (= ?:Lexeme:Literal:Natural:Value (ConvertBinaryTokenStringToNatural (. format) ? 1 0))
        (= ?:Lexeme:Literal:Natural:Format (LiteralFormat:MakeCompactNaturalLiteralFormat format))
        );;
     )local
     (GotoOperandField ?)
      >>
    
    #token CHARACTER [CHARACTER] "<squoted_character>"
      <<  (= ?:Lexeme:Literal:Character:Value (TokenStringCharacter ? 2))
      (= ?:Lexeme:Literal:Character:Format (LiteralFormat:MakeCompactCharacterLiteralFormat 0 0)) ; nothing special about character
      (GotoOperandField ?)
      >>
    
    
    %%OperandField
    
    #skip "<hblanks>"
      << (GotoEOLComment ?) >>
    #ifnotoken
      << (GotoEOLComment ?) >>
    
    -- Tokens signalling switch to index register modes
    #token ','   "\,"
       <<(GotoRegisterField ?)>>
    #token '['   "\["
       <<(GotoRegisterField ?)>>
    
    -- Operators for arithmetic syntax
    #token '!!'  "\!\!"
    #token '!'   "\!"
    #token '##'  "\#\#"
    #token '#'   "\#"
    #token '&'   "\&"
    #token '('   "\("
    #token ')'   "\)"
    #token '*'   "\*"
    #token '+'   "\+"
    #token '-'   "\-"
    #token '/'   "\/"
    #token '//'   "\/\/"
    #token '<'   "\<"
    #token '<'   "\<" 
    #token '<<'  "\<\<"
    #token '<='  "\<\="
    #token '</'  "\<\/"
    #token '='   "\="
    #token '>'   "\>"
    #token '>'   "\>"
    #token '>='  "\>\="
    #token '>>'  "\>\>"
    #token '>/'  "\>\/"
    #token '\\'  "\\"
    #token '|'   "\|"
    #token '||'  "\|\|"
    
    #token NUMBER [NATURAL] "<decimal_number>"
      << (local [format LiteralFormat:NaturalLiteralFormat]
        (;; (= ?:Lexeme:Literal:Natural:Value (ConvertDecimalTokenStringToNatural (. format) ? 0 0))
        (= ?:Lexeme:Literal:Natural:Format (LiteralFormat:MakeCompactNaturalLiteralFormat format))
        );;
     )local
      >>
    
    #token NUMBER [NATURAL] "\$ <hexadecimal_digit>+"
      << (local [format LiteralFormat:NaturalLiteralFormat]
        (;; (= ?:Lexeme:Literal:Natural:Value (ConvertHexadecimalTokenStringToNatural (. format) ? 1 0))
        (= ?:Lexeme:Literal:Natural:Format (LiteralFormat:MakeCompactNaturalLiteralFormat format))
        );;
     )local
      >>
    
    #token NUMBER [NATURAL] "\% <binary_digit>+"
      << (local [format LiteralFormat:NaturalLiteralFormat]
        (;; (= ?:Lexeme:Literal:Natural:Value (ConvertBinaryTokenStringToNatural (. format) ? 1 0))
        (= ?:Lexeme:Literal:Natural:Format (LiteralFormat:MakeCompactNaturalLiteralFormat format))
        );;
     )local
      >>
    
    -- Notice that an apparent register is accepted as a label in this mode
    #token IDENTIFIER [STRING] "<identifier>"
      <<  (local (;; (= [TokenScan natural] 1) ; process all string characters
             (= [TokenLength natural] ?:TokenCharacterCount)=
             (= [TokenString (reference TokenBodyT)] (. ?:TokenCharacters))
             (= [Result (reference string)] (. ?:Lexeme:Literal:String:Value))
             [ThisCharacterCode natural]
             (define Ordinala #61)
             (define Ordinalf #66)
             (define OrdinalA #41)
             (define OrdinalF #46)
         );;
         (;; (= (@ Result) `') ; start with empty string
         (while (<= TokenScan TokenLength)
          (;;   (= ThisCharacterCode (coerce natural TokenString:TokenScan))  
            (+= TokenScan) ; bump past character
            (ifthen (>= ThisCharacterCode Ordinala)
               (-= ThisCharacterCode #20) ; fold to upper case
            )ifthen
            (= (@ Result) (append (@ Result) (coerce character ThisCharacterCode)))=
    
            );;
         )while
         );;
      )local
      (= ?:Lexeme:Literal:String:Format (LiteralFormat:MakeCompactStringLiteralFormat 0))  ; nothing interesting in string
      >>
    
    %%Register -- operand field for TFR, ANDCC, ORCC, EXG opcodes
    
    #skip "<hblanks>"
    #ifnotoken << (GotoRegisterField ?) >>
    
    %%RegisterField -- handles registers and indexing mode syntax
    -- In this mode, names that look like registers are recognized as registers
    
    #skip "<hblanks>"
      << (GotoEOLComment ?) >>
    #ifnotoken
      << (GotoEOLComment ?) >>
    
    #token '['   "\["
    #token ']'   "\]"
    #token '--'  "\-\-"
    #token '++'  "\+\+"
    
    #token 'A'      "[aA]"
    #token 'B'      "[bB]"
    #token 'CC'     "[cC][cC]"
    #token 'DP'     "[dD][pP] | [dD][pP][rR]" -- DPR shouldnt be needed, but found one instance
    #token 'D'      "[dD]"
    #token 'Z'      "[zZ]"
    
    -- Index register designations
    #token 'X'      "[xX]"
    #token 'Y'      "[yY]"
    #token 'U'      "[uU]"
    #token 'S'      "[sS]"
    #token 'PCR'    "[pP][cC][rR]"
    #token 'PC'     "[pP][cC]"
    
    #token ','    "\,"
    
    -- Operators for arithmetic syntax
    #token '!!'  "\!\!"
    #token '!'   "\!"
    #token '##'  "\#\#"
    #token '#'   "\#"
    #token '&'   "\&"
    #token '('   "\("
    #token ')'   "\)"
    #token '*'   "\*"
    #token '+'   "\+"
    #token '-'   "\-"
    #token '/'   "\/"
    #token '<'   "\<"
    #token '<'   "\<" 
    #token '<<'  "\<\<"
    #token '<='  "\<\="
    #token '<|'  "\<\|"
    #token '='   "\="
    #token '>'   "\>"
    #token '>'   "\>"
    #token '>='  "\>\="
    #token '>>'  "\>\>"
    #token '>|'  "\>\|"
    #token '\\'  "\\"
    #token '|'   "\|"
    #token '||'  "\|\|"
    
    #token NUMBER [NATURAL] "<decimal_number>"
      << (local [format LiteralFormat:NaturalLiteralFormat]
        (;; (= ?:Lexeme:Literal:Natural:Value (ConvertDecimalTokenStringToNatural (. format) ? 0 0))
        (= ?:Lexeme:Literal:Natural:Format (LiteralFormat:MakeCompactNaturalLiteralFormat format))
        );;
     )local
      >>
    
    ... [snip]
    
    %% -- end M6809.lex
    

    ****************分析器***

    -- M6809.ATG: Motorola 6809 assembly code parser
    -- (C) Copyright 1989;1999-2002 Ira D. Baxter; All Rights Reserved
    
    m6809 = sourcelines ;
    
    sourcelines = ;
    sourcelines = sourcelines sourceline EOL ;
      <<PrettyPrinter>>: { V(CV(sourcelines[1]),H(sourceline,A<eol>(EOL))); }
    
    -- leading opcode field symbol should be treated as keyword.
    
    sourceline = ;
    sourceline = labels ;
    sourceline = optional_labels 'EQU' expression ;
      <<PrettyPrinter>>: { H(optional_labels,A<opcode>('EQU'),A<operand>(expression)); }
    sourceline = LABEL 'SET' expression ;
      <<PrettyPrinter>>: { H(A<firstlabel>(LABEL),A<opcode>('SET'),A<operand>(expression)); }
    sourceline = optional_label instruction ;
      <<PrettyPrinter>>: { H(optional_label,instruction); }
    sourceline = optional_label optlabelleddirective ;
      <<PrettyPrinter>>: { H(optional_label,optlabelleddirective); }
    sourceline = optional_label implicitdatadirective ;
      <<PrettyPrinter>>: { H(optional_label,implicitdatadirective); }
    sourceline = unlabelleddirective ;
    sourceline = '?ERROR' ;
      <<PrettyPrinter>>: { A<opcode>('?ERROR'); }
    
    optional_label = labels ;
    optional_label = LABEL ':' ;
      <<PrettyPrinter>>: { H(A<firstlabel>(LABEL),':'); }
    optional_label = ;
    
    optional_labels = ;
    optional_labels = labels ;
    labels = LABEL ;
      <<PrettyPrinter>>: { A<firstlabel>(LABEL); }
    labels = labels ',' LABEL ;
      <<PrettyPrinter>>: { H(labels[1],',',A<otherlabels>(LABEL)); }
    
    unlabelleddirective = 'END' ;
      <<PrettyPrinter>>: { A<opcode>('END'); }
    unlabelleddirective = 'END' expression ;
      <<PrettyPrinter>>: { H(A<opcode>('END'),A<operand>(expression)); }
    unlabelleddirective = 'IF' expression EOL conditional ;
      <<PrettyPrinter>>: { V(H(A<opcode>('IF'),H(A<operand>(expression),A<eol>(EOL))),CV(conditional)); }
    unlabelleddirective = 'IFDEF' IDENTIFIER EOL conditional ;
      <<PrettyPrinter>>: { V(H(A<opcode>('IFDEF'),H(A<operand>(IDENTIFIER),A<eol>(EOL))),CV(conditional)); }
    unlabelleddirective = 'IFUND' IDENTIFIER EOL conditional ;
      <<PrettyPrinter>>: { V(H(A<opcode>('IFUND'),H(A<operand>(IDENTIFIER),A<eol>(EOL))),CV(conditional)); }
    unlabelleddirective = 'INCLUDE' FILENAME ;
      <<PrettyPrinter>>: { H(A<opcode>('INCLUDE'),A<operand>(FILENAME)); }
    unlabelleddirective = 'LIST' expression ;
      <<PrettyPrinter>>: { H(A<opcode>('LIST'),A<operand>(expression)); }
    unlabelleddirective = 'NAME' IDENTIFIER ;
      <<PrettyPrinter>>: { H(A<opcode>('NAME'),A<operand>(IDENTIFIER)); }
    unlabelleddirective = 'ORG' expression ;
      <<PrettyPrinter>>: { H(A<opcode>('ORG'),A<operand>(expression)); }
    unlabelleddirective = 'PAGE' ;
      <<PrettyPrinter>>: { A<opcode>('PAGE'); }
    unlabelleddirective = 'PAGE' HEADING ;
      <<PrettyPrinter>>: { H(A<opcode>('PAGE'),A<operand>(HEADING)); }
    unlabelleddirective = 'PCA' expression ;
      <<PrettyPrinter>>: { H(A<opcode>('PCA'),A<operand>(expression)); }
    unlabelleddirective = 'PCC' expression ;
      <<PrettyPrinter>>: { H(A<opcode>('PCC'),A<operand>(expression)); }
    unlabelleddirective = 'PSR' expression ;
      <<PrettyPrinter>>: { H(A<opcode>('PSR'),A<operand>(expression)); }
    unlabelleddirective = 'TABS' numberlist ;
      <<PrettyPrinter>>: { H(A<opcode>('TABS'),A<operand>(numberlist)); }
    unlabelleddirective = 'TITLE' HEADING ;
      <<PrettyPrinter>>: { H(A<opcode>('TITLE'),A<operand>(HEADING)); }
    unlabelleddirective = 'WITH' settings ;
      <<PrettyPrinter>>: { H(A<opcode>('WITH'),A<operand>(settings)); }
    
    settings = setting ;
    settings = settings ',' setting ;
      <<PrettyPrinter>>: { H*; }
    setting = 'WI' '=' NUMBER ;
      <<PrettyPrinter>>: { H*; }
    setting = 'DE' '=' NUMBER ;
      <<PrettyPrinter>>: { H*; }
    setting = 'M6800' ;
    setting = 'M6801' ;
    setting = 'M6809' ;
    setting = 'M6811' ;
    
    -- collects lines of conditional code into blocks
    conditional = 'ELSEIF' expression EOL conditional ;
      <<PrettyPrinter>>: { V(H(A<opcode>('ELSEIF'),H(A<operand>(expression),A<eol>(EOL))),CV(conditional[1])); }
    conditional = 'ELSE' EOL else ;
      <<PrettyPrinter>>: { V(H(A<opcode>('ELSE'),A<eol>(EOL)),CV(else)); }
    conditional = 'FIN' ;
      <<PrettyPrinter>>: { A<opcode>('FIN'); }
    conditional = sourceline EOL conditional ;
      <<PrettyPrinter>>: { V(H(sourceline,A<eol>(EOL)),CV(conditional[1])); }
    
    else = 'FIN' ;
      <<PrettyPrinter>>: { A<opcode>('FIN'); }
    else = sourceline EOL else ;
      <<PrettyPrinter>>: { V(H(sourceline,A<eol>(EOL)),CV(else[1])); }
    
    -- keyword-less directive, generates data tables
    
    implicitdatadirective = implicitdatadirective ',' implicitdataitem ;
      <<PrettyPrinter>>: { H*; }
    implicitdatadirective = implicitdataitem ;
    
    implicitdataitem = '#' expression ;
      <<PrettyPrinter>>: { A<operand>(H('#',expression)); }
    implicitdataitem = '+' expression ;
      <<PrettyPrinter>>: { A<operand>(H('+',expression)); }
    implicitdataitem = '-' expression ;
      <<PrettyPrinter>>: { A<operand>(H('-',expression)); }
    implicitdataitem = expression ;
      <<PrettyPrinter>>: { A<operand>(expression); }
    implicitdataitem = STRING ;
      <<PrettyPrinter>>: { A<operand>(STRING); }
    
    -- instructions valid for m680C (see Software Dynamics ASM manual)
    instruction = 'ABA' ;
      <<PrettyPrinter>>: { A<opcode>('ABA'); }
    instruction = 'ABX' ;
      <<PrettyPrinter>>: { A<opcode>('ABX'); }
    
    instruction = 'ADC' 'A' operandfetch ;
      <<PrettyPrinter>>: { H(A<opcode>(H('ADC','A')),A<operand>(operandfetch)); }
    instruction = 'ADC' 'B' operandfetch ;
      <<PrettyPrinter>>: { H(A<opcode>(H('ADC','B')),A<operand>(operandfetch)); }
    instruction = 'ADCA' operandfetch ;
      <<PrettyPrinter>>: { H(A<opcode>('ADCA'),A<operand>(operandfetch)); }
    instruction = 'ADCB' operandfetch ;
      <<PrettyPrinter>>: { H(A<opcode>('ADCB'),A<operand>(operandfetch)); }
    instruction = 'ADCD' operandfetch ;
      <<PrettyPrinter>>: { H(A<opcode>('ADCD'),A<operand>(operandfetch)); }
    
    instruction = 'ADD' 'A' operandfetch ;
      <<PrettyPrinter>>: { H(A<opcode>(H('ADD','A')),A<operand>(operandfetch)); }
    instruction = 'ADD' 'B' operandfetch ;
      <<PrettyPrinter>>: { H(A<opcode>(H('ADD','B')),A<operand>(operandfetch)); }
    instruction = 'ADDA' operandfetch ;
      <<PrettyPrinter>>: { H(A<opcode>('ADDA'),A<operand>(operandfetch)); }
    
    [..snip...]
    
    -- condition code mask for ANDCC and ORCC
    conditionmask = '#' expression ;
      <<PrettyPrinter>>: { H*; }
    conditionmask = expression ;
    
    target = expression ;
    
    operandfetch = '#' expression ; --immediate
      <<PrettyPrinter>>: { H*; }
    
    operandfetch = memoryreference ;
    
    operandstore = memoryreference ;
    
    memoryreference = '[' indexedreference ']' ;
      <<PrettyPrinter>>: { H*; }
    memoryreference = indexedreference ;
    
    indexedreference = offset ;
    indexedreference = offset ',' indexregister ;
      <<PrettyPrinter>>: { H*; }
    indexedreference = ',' indexregister ;
      <<PrettyPrinter>>: { H*; }
    indexedreference = ',' '--' indexregister ;
      <<PrettyPrinter>>: { H*; }
    indexedreference = ',' '-' indexregister ;
      <<PrettyPrinter>>: { H*; }
    indexedreference = ',' indexregister '++' ;
      <<PrettyPrinter>>: { H*; }
    indexedreference = ',' indexregister '+' ;
      <<PrettyPrinter>>: { H*; }
    
    offset = '>' expression ; -- page zero ref
      <<PrettyPrinter>>: { H*; }
    offset = '<' expression ; -- long reference
      <<PrettyPrinter>>: { H*; }
    offset = expression ;
    offset = 'A' ;
    offset = 'B' ;
    offset = 'D' ;
    
    registerlist = registername ;
    registerlist = registerlist ',' registername ;
      <<PrettyPrinter>>: { H*; }
    
    registername = 'A' ;
    registername = 'B' ;
    registername = 'CC' ;
    registername = 'DP' ;
    registername = 'D' ;
    registername = 'Z' ;
    registername = indexregister ;
    
    indexregister = 'X' ;
    indexregister = 'Y' ;
    indexregister = 'U' ;  -- not legal on M6811
    indexregister = 'S' ;
    indexregister = 'PCR' ;
    indexregister = 'PC' ;
    
    expression = sum '=' sum ;
      <<PrettyPrinter>>: { H*; }
    expression = sum '<<' sum ;
      <<PrettyPrinter>>: { H*; }
    expression = sum '</' sum ;
      <<PrettyPrinter>>: { H*; }
    expression = sum '<=' sum ;
      <<PrettyPrinter>>: { H*; }
    expression = sum '<' sum ;
      <<PrettyPrinter>>: { H*; }
    expression = sum '>>' sum ;
      <<PrettyPrinter>>: { H*; }
    expression = sum '>/' sum ;
      <<PrettyPrinter>>: { H*; }
    expression = sum '>=' sum ;
      <<PrettyPrinter>>: { H*; }
    expression = sum '>' sum ;
      <<PrettyPrinter>>: { H*; }
    expression = sum '#' sum ;
      <<PrettyPrinter>>: { H*; }
    expression = sum ;
    
    sum = product ;
    sum = sum '+' product ;
      <<PrettyPrinter>>: { H*; }
    sum = sum '-' product ;
      <<PrettyPrinter>>: { H*; }
    sum = sum '!' product ;
      <<PrettyPrinter>>: { H*; }
    sum = sum '!!' product ;
      <<PrettyPrinter>>: { H*; }
    
    product = term '*' product ;
      <<PrettyPrinter>>: { H*; }
    product = term '||' product ; -- wrong?
      <<PrettyPrinter>>: { H*; }
    product = term '/' product ;
      <<PrettyPrinter>>: { H*; }
    product = term '//' product ;
      <<PrettyPrinter>>: { H*; }
    product = term '&' product ;
      <<PrettyPrinter>>: { H*; }
    product = term '##' product ;
      <<PrettyPrinter>>: { H*; }
    product = term ;
    
    term = '+' term ;
      <<PrettyPrinter>>: { H*; }
    term = '-' term ; 
      <<PrettyPrinter>>: { H*; }
    term = '\\' term ; -- complement
      <<PrettyPrinter>>: { H*; }
    term = '&' term ; -- not
    
    term = IDENTIFIER ;
    term = NUMBER ;
    term = CHARACTER ;
    term = '*' ;
    term = '(' expression ')' ;
      <<PrettyPrinter>>: { H*; }
    
    numberlist = NUMBER ;
    numberlist = numberlist ',' NUMBER ;
      <<PrettyPrinter>>: { H*; }
    
        2
  •  3
  •   Greg Hewgill    17 年前

    BNF更一般地用于结构化的、嵌套的语言,如Pascal、C++,或者实际上来自Algol家族的任何东西(包括C语言等现代语言)。如果我正在实现一个汇编程序,我可能会使用一些简单的正则表达式来模式匹配操作码和操作数。我用了Z80汇编语言已经有一段时间了,但是你可能会用到如下的语言:

    /\s*(\w{2,3})\s+((\w+)(,\w+)?)?/
    

    这将匹配由两个或三个字母的操作码和一个或两个用逗号分隔的操作数组成的任何行。在提取这样的汇编行之后,您将看到操作码并为指令生成正确的字节,包括操作数的值(如果适用)。

    我在上面使用正则表达式概述的解析器类型将被称为“特殊”解析器,这实质上意味着您可以在某种块基础上(在汇编语言的情况下,通过文本行)拆分和检查输入。

        3
  •  2
  •   bobince    17 年前

    我认为你不需要过度考虑。没有必要让一个解析器把ld a、a分解成一个加载操作、目标和源寄存器,这时您可以将整个字符串(modulo case和whitespace)直接匹配到一个操作码中。

    没有那么多的操作码,它们的排列方式也不足以让您真正从分析和理解汇编程序IMO中获益。显然,您需要一个用于字节/地址/索引参数的分析器,但除此之外,我只需要一对一的查找。

    推荐文章